Vimarsana.com

Efficient LLM inference

• Source: artfintel.com
On quantization, distillation, and efficiency