Intermediate
Inference engines explained
Ollama, llama.cpp, vLLM, MLX, LM Studio. They all run models, but they are not the same thing. Here is what an inference engine does and how to pick one.
Quantisation, inference engines, model formats, and getting good performance.
Ollama, llama.cpp, vLLM, MLX, LM Studio. They all run models, but they are not the same thing. Here is what an inference engine does and how to pick one.
Quantisation is what makes local AI possible on consumer hardware. What the formats mean, what you give up, and how to choose one for your memory budget.