Skip to content

Local LLM

A local LLM runs inference on hardware you own, keeping data on premises instead of sending prompts to a cloud API.

  • on-premises LLM
  • llama.cpp
  • Ollama
  • vLLM
  • open weight models
  • self-hosted AI
  • inference server
  • VRAM sizing
  • private LLM deployment