Speed vs accuracy
[One sentence on the headline. Averaged across all four benchmarks. Up and to the left is better.]
The benchmarks
Accuracy
[One sentence per notable result. Where Pip wins, where it doesn’t.]
| Model | Invoices | MuSiQue | CORD | DocILE | Average |
|---|---|---|---|---|---|
| Pip | [X]% | [X]% | [X]% | [X]% | [X]% |
| Jev | [X]% | [X]% | [X]% | [X]% | [X]% |
| GPT 6.0 Astra | [X]% | [X]% | [X]% | [X]% | [X]% |
| Opus 5.5 | [X]% | [X]% | [X]% | [X]% | [X]% |
Speed
Wall-clock time to finish 1,000 items of the invoice workflow. Pip: the whole job in one request. Jev: questions sent in groups of [100] per request, the most its API accepts [confirm]. LLMs: one question at a time, in sequence.
Cost
Total spend for the full run across all four benchmarks, at each provider’s published rates on [date].
How we ran it
[Model versions and dates. Pip: one request per job. Jev: questions in groups of [100] per request. LLMs: one call per question, run in sequence. Settings: temperature, max tokens, how answers were parsed. What counted as correct.]
Limits
[Where this comparison is weak. Rate limits vary by account tier. LLMs could be prompted differently. Name any benchmark where another model beat Pip.]