Technical notes on LLM inference, neural network quantization, on-device AI, and performance engineering.
Browse the notes