Efficient LLM inference January 4, 2024 • Source: artfintel.com On quantization, distillation, and efficiency