Research
What It Is
Measurement framework for multi-model orchestration using self-consistency variance routing. Evaluated on 1,510 tasks across 4 benchmarks with 7,550+ auditable runs.
The Problem
Most multi-model ensembles run full inference on every query regardless of difficulty. ACAR routes by measured complexity — skipping expensive ensembling when a single model is sufficient.
Approach
- Self-consistency variance as a proxy for query complexity
- Threshold-based routing: single model vs full ensemble dispatch
- Evaluated across MMLU, HotpotQA, MedQA, and LegalBench benchmarks
- Immutable experiment artifacts with full audit logs across 7,550+ runs
Results
- 55.6% routing accuracy across 1,510 evaluation tasks
- Avoided full ensembling on 54.2% of tasks — significant cost reduction
- Evaluated on 4 benchmarks with 7,550+ auditable runs
Stack
PythonOpenAI APIAnthropic APIStatistical evaluationLaTeX