Deepinfra
A strong candidate when its available model produces reliable results on your evaluation set.
Works well
- Can be assessed with the same task-specific prompts and scoring rubric used for Together AI.
- Lets an existing Deepinfra user test quality without changing the application first.
- May meet the task's quality threshold even if another provider leads on an unrelated benchmark.
Trade-offs
- A favorable result for one model does not establish quality for the rest of the catalog.
- Differences in model versions or default settings can invalidate a direct comparison.