Skip to content

Our local LLMs in action

These are the local open-source models we currently rely on – all running on our own infrastructure, in our office basement.

Reasoning & coding

The thinkers: complex analysis, business logic and help with programming.

deepseek-v4:flash
Flagship: analysis, booking logic, complex coding
DeepSeek AI (CN) ~102 GB RAM MIT
kimi-linear:48b
Everyday sprinter, long documents & context understanding
Moonshot AI (CN) 30 GB RAM MIT
qwen3-coder:30b
Coding specialist for fast iterations & agent tasks
Alibaba / Qwen (CN) 18 GB RAM Apache 2.0
qwen3.6:35b
Mid-range reserve
Alibaba / Qwen (CN) 23 GB RAM Apache 2.0

Vision & documents

The eyes: understanding images and turning scans into usable text.

qwen3-vl:32b
Detailed image understanding: screenshots, diagrams, receipts
Alibaba / Qwen (CN) 20 GB RAM Apache 2.0
qwen3-vl:8b
Fast vision variant for simple image tasks
Alibaba / Qwen (CN) 6,1 GB RAM Apache 2.0
gemma4:31b
Image input + polished writing
Google 19 GB RAM Gemma license
deepseek-ocr
PDF/scan → Markdown
DeepSeek AI (CN) 6,7 GB RAM MIT

Retrieval & search

The memory: finds the right content in large data sets.

qwen3-embedding:8b
Vectors for similarity search
Alibaba / Qwen (CN) 4,7 GB RAM Apache 2.0

Our infrastructure

This is where our local models run.

Hardware
MacBook Pro 16" M5 Max, 128 GB RAM.
Model management
LiteLLM for management and authorization, Ollama runs the models.
Access & security
Cloudflare Tunnel to the local MacBook.
Power
UPS Ubiquiti UniFi.

Usage

How intensively our models are currently in use.

73.4M
Tokens this month
Input 72.6M
Output 758'548
3'170
Requests this month

Archive

Models we've replaced – for the record.

Reasoning & coding

Reasoning & coding
Old model Replaced by
qwen3.5:122b deepseek-v4:flash
qwen3.6:35b-a3b kimi-linear:48b
qwen3.6:27b qwen3.6:35b
qwen3.5:9b kimi-linear:48b
qwen3.5:4b kimi-linear:48b

Vision & documents

Vision & documents
Old model Replaced by
gemma4:12b qwen3-vl:8b
gemma4:e4b qwen3-vl:8b

Retrieval & search

Retrieval & search
Old model Replaced by
qwen3-embedding:4b qwen3-embedding:8b
qwen3-reranker:8b

Discover more