A 90-second walkthrough
Drop your AI traces in. Get back a ranked, evidence-backed optimization plan — then run the experiment and record the decision, all in the browser. Every number you're about to see is deterministic arithmetic on a real trace file. Nothing is invented, nothing is uploaded.
▮ parsing… normalizing 46 calls → 10 traces · pricing per call · checking outcomes
LLM formatting step "Markdown Formatter" consumes 25.3% of priced sample cost
Why this first: low risk, high confidence, affects 70% of traces, represents 25.3% of priced sample cost, and targets a specific auxiliary component that can be validated independently.
Assumption, stated: a deterministic formatter has ~zero LLM cost. An estimated maximum, never claimed as observed. The experiment below tests it.
Reduce answer-generation context
up to $0.1196 input-cost reduction (32.6%)
Upper-bound simulation. Retriever metadata in the same traces reports 20 of 175 retrieved documents actually used (11%). The scenario assumes context scales with retrieved documents — validate before relying on it.
A deterministic or template-based formatter can produce equivalent output for "Markdown Formatter" at negligible LLM cost.
primary metric: cost per task ↓ · guardrail: success rate ±2%
10 traces · 32 LLM calls · coverage 100%
cost/task $0.0367 · success 90%
Warning: baseline is synthetic and the variant is a user upload — flagged automatically, results may not be comparable.
$0.0367 → $0.0015 (−95.9%)
Passes — the variant meets the configured arithmetic success condition on this pair of snapshots. Two samples compared honestly; no statistical claims.
| guardrail | baseline | variant | change | allowed | result |
|---|---|---|---|---|---|
| Success rate | 90.0% | 90.0% | 0.0% | ±2% | Pass (unchanged) |
| Markdown validity | — | — | — | — | Manual check |
✓ Decision recorded: adopted · "validated on holdout traces" · experiment archived locally
Upload→Evidence→ Best first action→Experiment→ Comparison→Decision
Observability tools stop at "here's what you spent." Quorum turns traces into the next experiment worth running — with every figure labeled measured, calculated, or simulated, and every recommendation shipped with its own validation plan. Fully client-side. Four deterministic detectors. Zero LLM calls to produce a recommendation.
Step 1 of 9