Get API key

Blog

Writing on LLM API costs

Every post here starts from published vendor prices and ends with arithmetic you can reproduce. No benchmarks we did not run, no numbers we cannot show the source for.

· 9 min read

How to reduce your LLM API costs

Most advice about cutting inference spend is vague. This is the concrete version: where the money goes, which five levers move it, and how much each one is worth on a real monthly volume.

Read the post

· 7 min read

OpenAI vs Anthropic API pricing

Both vendors publish per-million-token rates, but they price context, output, and caching differently enough that the cheaper choice depends entirely on your workload shape.

Read the post

· 6 min read

GPT-5.5 vs Claude Sonnet 4.6, priced honestly

Two frontier models from two vendors, compared on the only dimension we can verify: what they cost to run at published rates on identical workloads.

Read the post

· 7 min read

How to estimate your monthly LLM API cost

Forecasting an inference bill is arithmetic, not prophecy. Here is the method: four numbers to log, one formula, and the mistakes that make estimates come in low.

Read the post

· 6 min read

What prompt caching is actually worth

Cached input is billed at roughly a tenth of the normal input rate. Here is what qualifies, what silently does not, and how to structure prompts so the discount applies.

Read the post