Against a cached prompt
The honest comparison. Few-shot examples are a fixed prefix, so they cost a tenth, and the fine-tune's inference premium never earns the training run back.
- model
- gpt-4o-mini
- examples
- 500
- tokens-per-example
- 800
- epochs
- 3
- prompt-tokens
- 400
- few-shot-tokens
- 3000
- output-tokens
- 200
- cache-few-shot
- yes
- requests-per-month
- 100000
- hosting-per-hour
- 0