Usage planning

How to estimate deepinfra pricing for your workload

Start with a model, an estimated number of requests, and the tokens each request may use. Deepinfra helps you reason through those inputs, but this page does not display a live rate or produce an official quote.

Deepinfra model and inference visual

prerequisites

Gather workload measurements first. A model name alone cannot tell you what repeated requests will consume.

Required Optional
  • Choose the exact model you intend to use, rather than estimating from a model family. — Different model variants can have different usage units and rates.

  • Record the current published rate and billing units for that model at the point of use. — Do not treat an example on this page as a current rate.

  • Estimate requests over a defined period and measure typical input and output separately. — A long response can change the result even when the prompt stays short.

  • Keep a longer or more complex request as a second test case.optional — It helps show how sensitive your estimate is to workload changes.

one full run-through

Before doing the arithmetic, mark what the estimate cannot establish. Each gap calls for a measurement or a current source.

1

It cannot supply a live rate

A model's published rate may change, and this page does not fetch it. An estimate built from an old figure can be precise arithmetic with the wrong answer.

What to do instead

Look up the exact model's current rate and units immediately before calculating.

2

It cannot predict response length

Identical prompts can produce outputs of different lengths. A single short response is a weak basis for repeated-workload planning.

What to do instead

Sample representative tasks and include a longer-response case.

3

It cannot equate unlike units

Chat, embeddings, and other inference tasks may be measured differently. Multiplying every workload by one token assumption conceals that difference.

What to do instead

Make a separate estimate for each task and its published usage unit.

one full run-through

For a chat workload, carry one representative request from measurement to a period estimate. Replace every sample quantity with your own observations.

  1. 1

    Measure a representative request

    Take a typical prompt and response from the task you actually expect to run. Record input tokens and output tokens separately; include instructions or other context sent with every request.

  2. 2

    Apply the model's units

    Find the current rate for the exact model and note whether input and output have separate rates. Convert each measured token count to the published unit, multiply by its matching rate, then add the results.

  3. 3

    Scale and check the range

    Multiply the per-request result by expected requests in the period. Repeat with a longer response and a heavier-use period. Treat the gap between results as a planning range, not a guaranteed total.

options table

  • MEASURE INPUT
  • MEASURE OUTPUT
  • CHECK CURRENT UNITS

Compare workloads on the same basis

Use three columns in your own comparison: option, usage unit, and measured volume. For a short chat task, compare its measured input and output against the chosen chat model's current units. For a long-context chat task, repeat the measurement with the full context you will actually send. For embeddings, use the embedding model's published unit instead of carrying over the chat calculation.

Keep the period and request count consistent across options. If one option changes both the model and the workload, label both changes; otherwise the comparison will not show which one drove the difference. Deepinfra presents this as a method for checking assumptions, not a table of live model rates.

what fails

A copied rate, a guessed response length, or a mixed-unit calculation can make an otherwise careful estimate misleading. Check the current model details, measure a real task, and recalculate when your workload changes. You can then explore an inference option with a clearer idea of which assumptions still need testing.

Make the first estimate useful, not falsely exact

  • Use the exact model variant
  • Keep input and output separate
  • Test a heavier-use case
Explore inference options

its own FAQ

The model and its published usage units are essential inputs to an estimate. Check the exact variant's current details rather than assuming that every model in a family uses the same rate.

Measure input and output for a representative request, apply the chosen model's current units and rates to each, then scale by your expected request count. Repeat with a longer response to see how much the result could move.

No. First check the embedding model's published usage unit, then measure the text you intend to process in that unit. Keep the embedding calculation separate from chat requests.

Requests can vary in context length, response length, and frequency. An estimate can also become stale if the rate or model choice changes, so revisit its inputs against observed work.

Try AI models
Try AI models