Skip to content

feat(cost): add estimate-cost command for budgeting runs before starting them - #23

Merged
malteos merged 4 commits into
mainfrom
feat/cost-estimation
Oct 2, 2026
Merged

malteos merged 4 commits into
mainfrom
feat/cost-estimation

Conversation

@malteos

@malteos malteos commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Supersedes #12.

Why

#12 estimated LLM token costs, and #18 worked out Google Translate's per-character cost by hand. Both answer the same question: what will this run cost before we start it? This PR answers it in one place for any billing model: per-token LLM APIs, per-character APIs, and per-hour hardware for self-hosted models.

What

New command commonlid estimate-cost. It keeps usage (what a provider meters) separate from rate cards (what it charges per unit), and reports every figure as expected (low-high).

commonlid estimate-cost --model GoogleTranslate-v2 --dataset commonlid
commonlid estimate-cost --model dspy:openai/gpt-5 --dataset commonlid
commonlid estimate-cost --model GlotLID --dataset commonlid --hardware aws:g5.xlarge --throughput 500:2000:5000
dspy_openai_gpt-5 on commonlid_nano (1,507 samples)
Rate card: LiteLLM model map (litellm 1.83.0)

meter             source   per sample     total                   rate     cost (expected, low-high)
input_tokens      counted  298            449,693                 $1.25/M  $0.56
output_tokens     assumed  16 (12-32)     24,112 (18,084-48,224)  $10/M    $0.24 ($0.18-$0.48)
reasoning_tokens  assumed  256 (0-1,024)  385,792 (0-1,543,168)   $10/M    $3.86 ($0.00-$15.43)

Total: $4.66 ($0.74-$16.48)

Each meter says where its number came from:

  • counted: exact, counted offline over every sample (characters sent to Google, prompt tokens via the model's tokenizer). No API calls.
  • assumed: a per-sample range for what can't be counted offline (output and reasoning tokens, hardware throughput). Override it with --assume METER=LOW:EXPECTED:HIGH or --throughput.
  • calibrated: optional --calibrate N predicts N random samples for real and replaces the assumptions with the measured mean and a 95% confidence interval. This costs money on paid APIs.

Prices come from rate cards that carry an "as of" date and a source URL. LLM token prices come from LiteLLM's model map. Other paid APIs declare their price on the model class, and a few hardware options (EC2 g5/g6, HF Inference Endpoints) are built in. commonlid list-rate-cards lists both. --hourly-rate and --rate METER=USD cover anything else.

Deliberately out of scope: free tiers, volume discounts, prediction-cache awareness (the estimate is always for a run from scratch), and any budget guard on run.

How models opt in

LIDModel gets a rate_card attribute, next to supported_languages, plus optional hooks. All default to "not billed per call":

  • rate_card: the model's RateCard, set on the class for fixed API prices. DSPy LLMs set it per instance from the model name, the same way they already set model_id.
  • _estimate_usage(texts): counted usage of already-preprocessed texts
  • usage_assumptions(): per-sample ranges for uncountable meters
  • measure_usage(texts): a live run reporting the usage of each request, used by --calibrate

cost/rate_cards.py holds only what is shared: hardware cards (keyed by name, e.g. aws:g5.xlarge) and the LiteLLM lookup. A RateCard has no id of its own: a model's card is identified by its model id, hardware by its name. A new get_model_class() in the registry lets list-rate-cards read each model's rate_card without instantiating it.

Implemented for:

  • Google v2/v3: rate_card at $20/M characters. Usage is characters, stripped, skipping blanks and clipped exactly as each wrapper sends them.
  • DSPy LLMs: feat: add estimate-tokens command for offline LLM token & cost estimation #12's tokenizer logic moved into the model. Reasoning tokens are only assumed for models LiteLLM marks as reasoning models. Measured usage is read from DSPy's LM history, with reasoning tokens split out of completion tokens so they aren't billed twice.

Results on commonlid_nano

model estimate
GoogleTranslate-v2 / v3 $6.10
gpt-5 $4.66 ($0.74-$16.48)
gpt-4o-mini $0.08 ($0.08-$0.10)

Google counts 305,178 billable characters versus 305,431 in #18 (0.08% fewer). Counting what the wrapper sends after stripping whitespace probably explains the gap.

Testing

  • make check passes: 350 tests, 96% coverage, ruff and mypy strict clean.
  • Ran live, offline: Google, gpt-5 and gpt-4o-mini estimates; hardware estimates with --throughput; and --calibrate 2000 on cld2.
  • Calibration runs one untimed warm-up batch first. The first batch ran about 5× slower and pulled the confidence interval down to $0.
  • Not run live: DSPy --calibrate against a paid API. It is covered by mocked tests.

…ing them

Separates usage (what a provider meters) from rate cards (what it charges)
so one estimator covers per-token LLM APIs, per-character Google Translate
and per-hour hardware. Every figure is a low/expected/high range; meters are
counted offline, assumed, or calibrated with an optional paid sample run.
Google Translate's rate cards move from the shared rate_cards module into
GoogleTranslateV2Model/V3Model as a `pricing` ClassVar, next to
`supported_languages`. rate_cards.py keeps only what is shared (hardware
cards, the LiteLLM lookup), so get_rate_card becomes get_hardware_card.
list-rate-cards finds model pricing through the new get_model_class.
Consistent with RateCard and the rest of the cost API. The rate_card()
method goes away: rate_card is a plain attribute, set on the class for
fixed API prices and per instance by DSPyLLMModel, as model_id already is.
It duplicated model_id for cards attached to a model. Hardware is still
named, by its HARDWARE_CARDS key: estimate_cost takes that name and records
it on CostEstimate.hardware. list-rate-cards lists model cards by model id.
@malteos
malteos merged commit a011fc6 into main Oct 2, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant