We Told a 70B LLM to Protect Your Wallet. It Started Refusing Legitimate Purchases.
A 50-scenario benchmark of decision models for fiduciary shopping agents — Clef-omni vs Jev vs Strands Decider, and why prompting alone backfires.
The Structural Defect in Modern Shopping AI
By Peng Jiang, TimoBuy — October 2026
Most AI shopping assistants have the same structural defect: they want you to buy things. E-commerce bots optimize for GMV. Frontier LLMs trained on web review hype mirror that enthusiasm. Ask "should I upgrade my M1 Pro to M3 Pro for web development?" and you get 3nm benchmarks and a checkout link — even though the workload never saturates the M1's memory bandwidth.
I build TimoBuy, a shopping assistant. I wanted it to act in the buyer's interest, which in the physical world usually means KEEP, WAIT, REPAIR, or SKIP — not BUY. So I benchmarked whether small specialized decision models actually do this better than frontier LLMs. Then I open-sourced everything.
Open-source repo, dataset, and evaluator (MIT, zero dependencies): https://github.com/roc-chiang/timo-decision-benchmark
Live playground: https://timobuy.com/demo
The Benchmark Methodology and Results
50 high-stakes consumer dilemmas across laptops, audio gear, espresso machines, ergonomic furniture, cameras, and phones. Each scenario has a ground-truth verdict labeled against a published methodology (datasets/LABELING_GUIDE.md in the repo). n=50, Wilson 95% CI [86.5%, 98.9%] — small sample, reported honestly.
Four architectures were evaluated across accuracy, fiduciary restraint, false negative rate, latency, and cost:
1. Rule-based router (retail baseline): 38.0% accuracy | 6.1% fiduciary restraint | 0.0% false negatives | 4.2 ms median latency | $0.01 per 1k decisions.
2. Frontier 70B+ (raw baseline): 76.0% accuracy | 63.6% fiduciary restraint | 0.0% false negatives | 820 ms median latency | $15.00 per 1k decisions.
3. Frontier 70B+ (fiduciary system prompt): 88.0% accuracy | 93.9% fiduciary restraint | 23.5% false negatives | 840 ms median latency | $15.50 per 1k decisions.
4. Timo decoupled decision architecture: 96.0% accuracy | 93.9% fiduciary restraint | 0.0% false negatives | 48 ms median latency | $0.45 per 1k decisions.
*Fiduciary restraint: share of non-essential purchase cases where the system correctly advised KEEP / WAIT / REPAIR / SKIP instead of pushing a purchase.
**False negatives: cases with genuine physical blockers where the system wrongly advised NOT to buy.
Finding 1: The Decision-Model Bake-Off
Before comparing against LLMs, I tested whether specialized decision heads are even viable for this job. Three candidates were evaluated:
Clef-omni (1.8B): 42ms P50 / 68ms P99 latency | $0.38 per 1k queries | 94.0% (47/50) constraint adherence | 0.0% over-correction rate.
Jev (0.8B): 18ms P50 / 31ms P99 latency | $0.12 per 1k queries | 84.0% (42/50) constraint adherence | 11.8% over-correction rate.
Strands Decider (3.2B): 85ms P50 / 140ms P99 latency | $0.85 per 1k queries | 96.0% (48/50) constraint adherence | 0.0% over-correction rate.
Jev is blistering fast but fumbles multi-hop mechanical reasoning (it once confused espresso extraction channeling with pump failure). Strands has the best raw fidelity but 2.2x the hosting footprint. Clef-omni was the sweet spot for a serverless edge envelope: no false negatives, 42ms latency, and $0.38 per thousand decisions. Full breakdown is documented in docs/decision_model_bakeoff.md in the repository.
Finding 2: The Ablation That Backfired
This is the part I didn't expect.
The obvious cheap fix — before building any architecture — is a strict fiduciary system prompt on the 70B model: "you are a fiduciary advisor, protect the user's money at all costs." I ran it as an ablation. Restraint jumped from 63.6% to 93.9%. It stopped hyping $1,200 audiophile cables.
Then it started refusing legitimate purchases. False negatives hit 23.5% (4/17):
Case TB-002: A user fine-tuning 70B models locally on a 16GB Intel laptop — a genuine OOM blocker. The prompted model advised KEEP and suggested "just use the free tier of Colab," ignoring the user's explicit local/offline constraint.
Case TB-026: An Ironman 70.3 athlete whose watch dies at 4.5 hours. The model advised KEEP and suggested "pausing the watch at water stations."
Tell a generalist model to be frugal and it doesn't become discerning — it becomes austere. It learns the direction (say no more often) without the discrimination (knowing which no's are wrong). That's what the decoupled architecture is for: the physical constraint engine checks whether a real blocker exists before the verdict stage is allowed to say KEEP. Same 93.9% restraint, zero false negatives.
Prompting buys restraint. It doesn't buy judgment.
Finding 3: The Economics Were Never Close
Even if the prompted 70B had worked, $15.50 per 1,000 decisions kills it as a router. A shopping assistant runs the decision on every query — at any real volume, the frontier model costs 34x the decision head while being 17x slower. The router was never a quality debate alone; it was always an economic one.
What We Shipped and Live Playground
The architecture is live in TimoBuy's decision playground: https://timobuy.com/demo
The playground is pre-loaded with the real cases from the benchmark, alongside live toggles between the raw 70B, prompted 70B, and the decoupled engine. You can inspect the intermediate constraint extractions in real time.
Reproducibility, Code, and Caveats
Everything is open-sourced under the MIT license at https://github.com/roc-chiang/timo-decision-benchmark — zero dependencies, clean Python evaluator, and the full 50-scenario dataset with annotated failure modes.
A transparent caveat: 50 scenarios is a high-signal benchmark for architectural direction, not a definitive statistical census across the entire global retail catalog. The Wilson 95% confidence interval is [86.5%, 98.9%]. We report these numbers openly so the community can extend the test harness to larger consumer categories.