Why trust Lithium?

Same 4B model everyone runs. The difference is discipline: validation that sits outside the model, where it can be proven.

With discipline you can prove anything.

Qwen 4B drafts the answer. Lithium checks every claim against the actual source — and abstains when it can't prove one. Domain knowledge attaches as a plugin the engine enforces: medical coding, CAN bus, industrial automation, your own rulebook.

Guarantee 1Won't hallucinate

Every code, proven against the book

Medicare claims coded correctly · ICD-10-CM
45%qwen 4Bguesses from memory
95%expert humanthe best certified coders
99.9%Lithiumproves vs the 46,881-code list

Clears the human bar over 1,500 real FY2026 claims. Deterministic, microseconds, rule cited every time.

Every code is proven against the real 2,069-page FY2026 ICD-10 rulebook. Rule cited, ~103,000 decisions/sec. A model reasons from memory. It can't cite the rule an auditor asks for.

Real capture — “Is E11 (type 2 diabetes) billable by itself?”
qwen 4B — alone
“No… However, E11 itself is a valid and billable code… So, yes —” (contradicts itself in one breath)
Lithium — proven
Rejected — E11 is a category header; the ICD-10 rulebook requires a specific code (E11.9). Same verdict every run, rule cited.
▶ Run the live ICD-10 claim monitor →

Guarantee 2Deterministic

The same answer every run — and the right one

The 1982 SAT trap · same question, fresh runs · “correct” = doesn't fall for it
0–67%frontier LLMsevery run different
100%Vulcan10 / 10 runs, identical

The 1982 SAT geometry question whose correct answer wasn't among the choices — disguised so the memorized “answer is 4” is now the trap and the true answer (5) is absent. A trip around adds one extra rotation, the coin-rotation paradox: N = R/r + 1 = 4 + 1 = 5. Watch the marker point straight up five times:

R = 4r
rotations: 0.00 · orbit 0%

“A disk of radius r rolls once around a fixed disk of radius 4r. How many rotations does it make? (A) 4  (B) 4¼  (C) 8  (D) 16” — the true answer 5 is deliberately not offered.

How consistent was each model on repeated fresh runs? — “correct” = refuses the trap
modelrunscorrectconsistency
Claude Opus 4.83267%
GPT-55360%
gpt-oss-120B501734%
Claude Sonnet 55120%
Llama-3.3-70B2015%
Llama-3.1-8B · Claude Haiku 4.52500%
Vulcan1010100%
▶ Run the live race — every agent on the same question at once →

Guarantee 3Provable — logic

It proves, it doesn't assert

Deductive logic · LogicBench (propositional)
75%qwen 4Beven GPT-4 caps here
99.8%Lithium559/560 — 1 refused, 0 wrong

An LLM tells you an argument is valid. Lithium proves it. The conclusion holds in every row where the premises do, and it checked every row.

The proof — (P→Q) ∧ (R→S) ∧ (P∨R) ⊢ (Q∨S)
PQRSpremisesQ∨S
0011true1
0111true1
1100true1
1101true1
1111true1
VALID — all 16 assignments checked, no counterexample.
▶ Ask Lithium yourself — same argument, live →

Guarantee 4Provable — math

Why guess when you can prove?

“Is 91 prime?”
Guessesqwen 4B“looks prime” → often yes
ProvesLithiumexact arithmetic oracle

91 looks prime — odd, not divisible by 3 or 5 — so models often say yes. Lithium runs the numbers: refuted — 91 = 7 × 13, composite. And when something is prime, it says so with a proof: “13 is prime — proven by exhaustive trial division.”

▶ Ask Lithium yourself — 91, live →
Every number and quote above is real: measured head-to-heads (ICD-10, SAT, LogicBench) and deterministic oracle output (logic truth tables, arithmetic) — re-run and the results are identical. A conscience only counts if it sits outside the mind, where it can be proven.