LUMIX-4B · APACHE 2.0 · CLEAN · ALIGNED

Lumix-4B

LUM · illumination  —  IX · integrated intelligence

A four-billion-parameter model purpose-built for Shadow. Meticulously trained to drive the agent loop like a model many times its size.

A 4B that feels like a 70B.

MODEL .Lumix-4B
PARAMS .4.0B-dense
MODALITY .vision+text
CONTEXT .262144
LICENSE .Apache-2.0
STATUS .pre-release

Local · llama.cpp · Q5_K_M · ~3 GB served

SHADOW DRIVABILITY HARNESS [ EVAL · MEASURED ]
100% tool-call JSON validity
0 malformed calls
CONTEXT WINDOW262,144 (256K)
FULL RUN · 10 TASKS~26s
DENSE · VISION-CAPABLE4B · ~3 GB
Local: Lumix-4B (clean) · llama.cpp · Q5_K_M

lumix why

Small model. Big-model behavior.

Where it counts for an agent — driving tools without breaking — Lumix behaves like a model many times its size. In Shadow's drivability harness it produced zero malformed tool calls and finished faster than every larger model it was tested beside.

01

Flawless tool-driving

100% valid tool-call JSON across the agentic suite — the thing that makes or breaks a small model as an agent.

02

Fastest of its fleet

A full 10-task agentic run in ~26s — quicker than the 35B and 80B it was benchmarked against.

03

Built for the loop

Trained specifically to drive Shadow — plan mode, tools, MCP, sub-agents. It knows the harness it lives in.

Head-to-head benchmarks vs. external models are in progress — coming soon.

lumix purpose

PURPOSE-BUILT

Shadow runs any model. This one was made for it.

Most models are trained to chat. Lumix was trained to drive — to plan, call tools, recover from a failed edit, and finish the task inside Shadow's agent loop. Shadow is a lightweight, private CLI, so it deserves a model small enough to live on your box and sharp enough to feel frontier-class. That's Lumix: illumination and integrated intelligence, tuned to the harness it runs in.

Tuned to the loop

Trained on how Shadow actually works — plan mode, tools, MCP, sub-agents — not generic chat transcripts.

Runs on your box

~3 GB served. A single consumer GPU — or good CPU — is enough. No cluster, no cloud bill.

Private by default

Local weights in Shadow's zero-telemetry CLI — your code and context never leave the machine.

lumix blend

The blend

We keep the exact recipe private — but we'll show you the shape of it. Lumix is a coding-first creator agent, reasoning-aware, with real finance/quant depth, and — unlike derisked models — safety and alignment kept in by design.

Coding57%
Safety & alignment (kept)18%
Reasoning13%
Finance & quant11%

Approximate composition of the training blend. The seed sources, teachers, and exact recipe stay in-house.

[ ILLUMINATED ]

It can see.

Lumix is a unified vision-language model — a creator agent that reads reference images, not just text. Point it at a screenshot, a mockup, a diagram; it works from what it sees.

[ ALIGNED ]

Clean by design.

Aligned on an aligned base — safety and refusals are a first-class part of the blend, kept in on purpose. The deliberate opposite of a derisked model: a creator agent you can put in front of anyone.

lumix specs

[ UNDER THE HOOD ]

Under the hood.

Parameters
4B dense
Modality
Unified vision-language — text + images
Context
262,144  (256K)
Formats
Q4_K_M · Q5_K_M · Q6_K · Q8_0 · f16 + vision projector
Serving
Q5_K_M · ~3 GB · llama.cpp
Alignment
Aligned — safety & refusals retained by design

lumix builds

Quantized builds

GGUF, multiple formats — run it as light or as sharp as your box allows. Self-hosted downloads land here at launch.

QuantSizeNotesStatus
Q4_K_M ~2.5 GB Lightest — runs on the most modest boxes SOON
Q5_K_MDEFAULT ~3 GB Serving default — the build the harness runs SOON
Q6_K ~3.5 GB Balanced — more fidelity, still modest SOON
Q8_0 ~4.3 GB Near-lossless — for when fidelity matters most SOON
f16 ~8 GB Full precision — the reference weights SOON

Every build ships with the vision projector. Hosted by Blackfrost — no third-party account. Self-hosted downloads land here at launch.

lumix bench

[ LEADERBOARD · DATA AT LAUNCH ]

Benchmarks drop at launch.

Head-to-head results against external models are being measured now. This table fills the moment the numbers are final — no provisional scores.

Model Params Score Δ vs Lumix Notes
Lumix-4B OURS 4B dense Apache-2.0 · vision · aligned

Methodology and full results published with the model. Data supplied at release.

lumix drive

[ SHADOW HARNESS · FULL RESULTS PENDING ]

Built to be driven. Full results land at release.

The drivability harness is being expanded. The headline stats live in the hero — the full exec-verified task suite publishes here at launch.

DRIVABILITY HARNESS · FULL RUN [ PENDING ]

MEASURED · 5|PENDING · 4

tool-call JSON validity 100% MEASURED
malformed tool calls 0 MEASURED
full run · 10 tasks ~26s MEASURED
context window 262,144 MEASURED
served footprint ~3GB MEASURED
task completion rate PENDING
avg per-task latency PENDING
token throughput PENDING
edge-case refusals PENDING
Full task suite, exec-verified. Data supplied at release.
✓ full results published at launch.

The light in the Shadow.

Lumix is the default local model for Shadow — a private, zero-telemetry CLI that runs any model you point it at.