Repository navigation
feat(cost): add estimate-cost command for budgeting runs before starting them - #23
Merged
Merged
Conversation
…ing them Separates usage (what a provider meters) from rate cards (what it charges) so one estimator covers per-token LLM APIs, per-character Google Translate and per-hour hardware. Every figure is a low/expected/high range; meters are counted offline, assumed, or calibrated with an optional paid sample run.
Google Translate's rate cards move from the shared rate_cards module into GoogleTranslateV2Model/V3Model as a `pricing` ClassVar, next to `supported_languages`. rate_cards.py keeps only what is shared (hardware cards, the LiteLLM lookup), so get_rate_card becomes get_hardware_card. list-rate-cards finds model pricing through the new get_model_class.
Consistent with RateCard and the rest of the cost API. The rate_card() method goes away: rate_card is a plain attribute, set on the class for fixed API prices and per instance by DSPyLLMModel, as model_id already is.
It duplicated model_id for cards attached to a model. Hardware is still named, by its HARDWARE_CARDS key: estimate_cost takes that name and records it on CostEstimate.hardware. list-rate-cards lists model cards by model id.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Supersedes #12.
Why
#12 estimated LLM token costs, and #18 worked out Google Translate's per-character cost by hand. Both answer the same question: what will this run cost before we start it? This PR answers it in one place for any billing model: per-token LLM APIs, per-character APIs, and per-hour hardware for self-hosted models.
What
New command
commonlid estimate-cost. It keeps usage (what a provider meters) separate from rate cards (what it charges per unit), and reports every figure as expected (low-high).Each meter says where its number came from:
--assume METER=LOW:EXPECTED:HIGHor--throughput.--calibrate Npredicts N random samples for real and replaces the assumptions with the measured mean and a 95% confidence interval. This costs money on paid APIs.Prices come from rate cards that carry an "as of" date and a source URL. LLM token prices come from LiteLLM's model map. Other paid APIs declare their price on the model class, and a few hardware options (EC2 g5/g6, HF Inference Endpoints) are built in.
commonlid list-rate-cardslists both.--hourly-rateand--rate METER=USDcover anything else.Deliberately out of scope: free tiers, volume discounts, prediction-cache awareness (the estimate is always for a run from scratch), and any budget guard on
run.How models opt in
LIDModelgets arate_cardattribute, next tosupported_languages, plus optional hooks. All default to "not billed per call":rate_card: the model'sRateCard, set on the class for fixed API prices. DSPy LLMs set it per instance from the model name, the same way they already setmodel_id._estimate_usage(texts): counted usage of already-preprocessed textsusage_assumptions(): per-sample ranges for uncountable metersmeasure_usage(texts): a live run reporting the usage of each request, used by--calibratecost/rate_cards.pyholds only what is shared: hardware cards (keyed by name, e.g.aws:g5.xlarge) and the LiteLLM lookup. ARateCardhas no id of its own: a model's card is identified by its model id, hardware by its name. A newget_model_class()in the registry letslist-rate-cardsread each model'srate_cardwithout instantiating it.Implemented for:
rate_cardat $20/M characters. Usage is characters, stripped, skipping blanks and clipped exactly as each wrapper sends them.Results on
commonlid_nanoGoogle counts 305,178 billable characters versus 305,431 in #18 (0.08% fewer). Counting what the wrapper sends after stripping whitespace probably explains the gap.
Testing
make checkpasses: 350 tests, 96% coverage, ruff and mypy strict clean.--throughput; and--calibrate 2000on cld2.--calibrateagainst a paid API. It is covered by mocked tests.