Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
43 changes: 43 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -152,6 +152,49 @@ The `--api-base`, `--api-version`, `--api-key`, `--azure-ad-token`,
`--llm-n-threads` flags are only consumed by `dspy:` models and ignored
when every model is a classical one.

## Estimating cost before a run

Paid APIs and cloud hardware cost real money, so `estimate-cost` budgets a
run up front. It reports every figure as **expected (low-high)**:

```bash
commonlid estimate-cost --model GoogleTranslate-v2 --dataset commonlid
commonlid estimate-cost --model dspy:openai/gpt-5 --dataset commonlid
commonlid estimate-cost --model GlotLID --dataset commonlid \
--hardware aws:g5.xlarge --throughput 500:2000:5000
```

```
dspy_openai_gpt-5 on commonlid_nano (1,507 samples)
Rate card: LiteLLM model map (litellm 1.83.0)

meter source per sample total rate cost (expected, low-high)
input_tokens counted 298 449,693 $1.25/M $0.56
output_tokens assumed 16 (12-32) 24,112 (18,084-48,224) $10/M $0.24 ($0.18-$0.48)
reasoning_tokens assumed 256 (0-1,024) 385,792 (0-1,543,168) $10/M $3.86 ($0.00-$15.43)

Total: $4.66 ($0.74-$16.48)
```

Each meter has a source, which tells you how much to trust it:

| source | meaning |
|---|---|
| `counted` | Exact, counted offline over every sample: characters sent to Google, prompt tokens via the model's tokenizer. No API calls. |
| `assumed` | A per-sample range for what can't be counted offline: generated and reasoning tokens, hardware throughput. Override it with `--assume METER=LOW:EXPECTED:HIGH` or `--throughput LOW:EXPECTED:HIGH`. |
| `calibrated` | Measured with `--calibrate N`, which predicts N random samples **for real** and replaces the assumptions with the mean and a 95% confidence interval. This costs money on paid APIs, and LLMs then need the same `--api-base`/auth flags as `run`. |

Prices come from rate cards with an "as of" date and a source URL. LLM token
prices come from LiteLLM's model map. Other paid APIs declare their price
on the model class (`rate_card`), and hardware prices are built in. `commonlid
list-rate-cards` lists both. Use `--hourly-rate USD` for hardware not
in that list, and `--rate METER=USD_PER_UNIT` to override any price. Rates
are linear list prices: free tiers and volume discounts are not modelled.
Prediction caches are ignored too, so the estimate is for a run from scratch.

Hardware throughput from `--calibrate` is measured on the machine running
the command, so it only applies to `--hardware` if that is the same hardware.

## Python API

The `commonlid` import auto-registers every shipped model and dataset, so
Expand Down
32 changes: 32 additions & 0 deletions docs/contributing/adding_a_model.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,6 +62,38 @@ pass a real confidence: a constant, or a number on an unbounded scale, is
worse than `None` because it reads like one. See the table in the README for
what each shipped model does.

### Reporting what a call costs

Models billed per call declare their price list and implement two optional
hooks, so that `commonlid estimate-cost` can budget a run before it starts:

```python
from commonlid.cost import Range, RateCard

rate_card = RateCard(
rates={"characters": 20.0 / 1_000_000}, # USD per unit
as_of="2026-10-02",
source="https://example.com/pricing",
)

def _estimate_usage(self, texts: Sequence[str]) -> dict[str, float]:
# Billable usage of these (already preprocessed) texts, counted offline.
return {"characters": float(sum(len(t) for t in texts))}

def usage_assumptions(self) -> dict[str, Range]:
# Per-sample usage that can't be counted offline, as low/expected/high.
return {"output_tokens": Range(12, 16, 32)}
```

When the price depends on the instance, as for LLMs priced by model name,
set `self.rate_card` in `__init__` instead.

To support `--calibrate`, override `measure_usage(texts)` as well. It should
predict `texts` for real and return the usage the API reported for each
request. Always give `rate_card` an `as_of` date and a source URL, since prices
drift. Local models need none of this: their cost
is compute time, which `--hardware` and `--throughput` cover.

### Adding model dependencies

If you are adding a model that requires additional dependencies, you can add them to the `pyproject.toml` file, under optional dependencies:
Expand Down
208 changes: 204 additions & 4 deletions src/commonlid/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@
from commonlid.core.registry import (
get_dataset,
get_model,
get_model_class,
list_datasets,
list_models,
)
Expand Down Expand Up @@ -163,22 +164,221 @@ def run(
).run()


def _resolve_model_spec(spec: str, llm_kwargs: dict[str, Any]) -> Any:
"""Resolve a CLI --model spec to a loaded :class:`LIDModel` instance."""
def _resolve_model_spec(
spec: str, llm_kwargs: dict[str, Any], *, require_api_base: bool = True
) -> Any:
"""Resolve a CLI --model spec to a loaded :class:`LIDModel` instance.

``require_api_base=False`` builds a DSPy model that is never called, as
for offline cost estimates.
"""
if spec.startswith(DSPY_SPEC_PREFIX):
llm_model_name = spec.removeprefix(DSPY_SPEC_PREFIX)
if not llm_model_name:
msg = "'dspy:' model spec requires a model name, e.g. 'dspy:azure/gpt-4o-mini'"
raise typer.BadParameter(msg)
if not llm_kwargs.get("api_base"):
if require_api_base and not llm_kwargs.get("api_base"):
msg = "DSPy LLM models require --api-base (e.g. your Azure endpoint URL)"
raise typer.BadParameter(msg)
from commonlid.models.dspy_llm import DSPyLLMModel

return DSPyLLMModel(llm_model_name=llm_model_name, **llm_kwargs)
return DSPyLLMModel(
llm_model_name=llm_model_name,
**{**llm_kwargs, "api_base": llm_kwargs["api_base"] or ""},
)
return get_model(spec)


def _parse_meter_options(values: list[str] | None, option: str, parse: Any) -> dict[str, Any]:
"""Parse repeated ``METER=VALUE`` options."""
parsed: dict[str, Any] = {}
for item in values or []:
meter, sep, value = item.partition("=")
if not sep or not meter:
msg = f"{option} expects METER=VALUE, got {item!r}"
raise typer.BadParameter(msg)
try:
parsed[meter] = parse(value)
except ValueError as exc:
msg = f"{option} {item!r}: {exc}"
raise typer.BadParameter(msg) from exc
return parsed


@app.command("estimate-cost")
def estimate_cost_cmd(
model: Annotated[
list[str],
typer.Option(
"--model",
"-m",
help="Model id or 'dspy:<llm-model-name>' (repeat to add more).",
),
],
dataset: Annotated[
list[str], typer.Option("--dataset", "-d", help="Dataset id (repeat to add more).")
],
assume: Annotated[
list[str] | None,
typer.Option(
"--assume",
help=(
"Per-sample usage for a meter, as METER=VALUE or METER=LOW:EXPECTED:HIGH "
"(e.g. reasoning_tokens=0:200:1000). Repeatable."
),
),
] = None,
rate: Annotated[
list[str] | None,
typer.Option(
"--rate",
help="Override the price of a meter, as METER=USD_PER_UNIT. Repeatable.",
),
] = None,
hardware: Annotated[
str | None,
typer.Option(
"--hardware",
help="Price compute time on this hardware card (see `list-rate-cards`).",
),
] = None,
hourly_rate: Annotated[
float | None,
typer.Option("--hourly-rate", help="Price compute time at this USD/hour."),
] = None,
throughput: Annotated[
str | None,
typer.Option(
"--throughput",
help="Samples/second on the hardware, as VALUE or LOW:EXPECTED:HIGH.",
),
] = None,
calibrate: Annotated[
int,
typer.Option(
"--calibrate",
help=(
"Predict N random samples for real and use the measured usage or "
"throughput instead of assumptions. Costs money on paid APIs."
),
),
] = 0,
seed: Annotated[int, typer.Option("--seed", help="Seed for --calibrate sampling.")] = 0,
batch_size: Annotated[int, typer.Option("--batch-size")] = 64,
as_json: Annotated[bool, typer.Option("--json", help="Output JSON instead of text.")] = False,
# --- DSPy LLM flags, only needed with --calibrate ---
api_base: Annotated[str | None, typer.Option("--api-base")] = None,
api_version: Annotated[str | None, typer.Option("--api-version")] = None,
api_key: Annotated[str | None, typer.Option("--api-key")] = None,
azure_ad_token: Annotated[bool, typer.Option("--azure-ad-token")] = False,
max_completion_tokens: Annotated[int | None, typer.Option("--max-completion-tokens")] = None,
llm_n_threads: Annotated[int, typer.Option("--llm-n-threads")] = 1,
) -> None:
"""Estimate what a run would cost, as a low / expected / high range.

Billable usage that can be counted offline (characters sent, prompt
tokens) is counted over every sample without calling any API. What
cannot (generated tokens, throughput) comes from per-model defaults,
--assume, or a small paid --calibrate run.
"""
from commonlid.cost import (
Range,
RateCard,
estimate_cost,
format_estimate,
get_hardware_card,
hourly_rate_card,
)

if hardware is not None and hourly_rate is not None:
msg = "pass either --hardware or --hourly-rate, not both"
raise typer.BadParameter(msg)
# Built-in hardware is passed by name so the estimate can say which it was.
hardware_spec: str | RateCard | None = hardware
if hourly_rate is not None:
hardware_spec = hourly_rate_card(hourly_rate)
try:
if hardware is not None:
get_hardware_card(hardware) # fail before counting a whole dataset
throughput_range = Range.parse(throughput) if throughput is not None else None
except (KeyError, ValueError) as exc:
raise typer.BadParameter(str(exc)) from exc
assumptions = _parse_meter_options(assume, "--assume", Range.parse)
rate_overrides = _parse_meter_options(rate, "--rate", float)

llm_kwargs = {
"api_base": api_base,
"api_version": api_version,
"api_key": api_key,
"azure_ad_token": azure_ad_token,
"max_completion_tokens": max_completion_tokens,
"n_threads": llm_n_threads,
"cache_dir": None,
}
models = [
_resolve_model_spec(spec, llm_kwargs, require_api_base=calibrate > 0) for spec in model
]
estimates = []
for dataset_id in dataset:
lid_dataset = get_dataset(dataset_id)
for lid_model in models:
try:
estimates.append(
estimate_cost(
lid_model,
lid_dataset,
assumptions=assumptions,
rate_overrides=rate_overrides,
hardware=hardware_spec,
throughput=throughput_range,
calibrate=calibrate,
seed=seed,
batch_size=batch_size,
)
)
except ValueError as exc:
raise typer.BadParameter(str(exc)) from exc

if as_json:
typer.echo(json.dumps([e.to_dict() for e in estimates], indent=2))
else:
typer.echo("\n\n".join(format_estimate(e) for e in estimates))


@app.command("list-rate-cards")
def list_rate_cards_cmd(
as_json: Annotated[bool, typer.Option("--json", help="Output JSON instead of text.")] = False,
) -> None:
"""List the built-in rate cards: each paid model's, then hardware.

LLM token prices are not listed; they come from LiteLLM's model map.
"""
from commonlid.cost import HARDWARE_CARDS

cards = {
model_id: card
for model_id in list_models()
if (card := get_model_class(model_id).rate_card) is not None
}
cards.update(HARDWARE_CARDS)
if as_json:
typer.echo(
json.dumps({
name: {
"rates": dict(card.rates),
"as_of": card.as_of,
"source": card.source,
"notes": list(card.notes),
}
for name, card in cards.items()
})
)
return
for name, card in cards.items():
rates = ", ".join(f"{meter}={usd:.4g}" for meter, usd in card.rates.items())
typer.echo(f"{name}: {rates} USD/unit ({'; '.join(card.notes)}, as of {card.as_of})")


@app.command()
def predict(
model: Annotated[str, typer.Option("--model", "-m", help="Model id.")],
Expand Down
2 changes: 2 additions & 0 deletions src/commonlid/core/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@
from commonlid.core.registry import (
get_dataset,
get_model,
get_model_class,
list_datasets,
list_models,
register_dataset,
Expand All @@ -17,6 +18,7 @@
"LIDPrediction",
"get_dataset",
"get_model",
"get_model_class",
"list_datasets",
"list_models",
"register_dataset",
Expand Down
Loading
Loading