← ToolsFree tool
GPU / vLLM TCO
Estimate when self-hosted GPU hours beat API list prices for a given throughput. Use it before committing to A100/H100 capacity for vLLM.
- 1Fill inputs
- 2Run
- 3Copy / export
Inputs
Required fields on the left
Results
Appears after you run
GPU vs API break-even
See whether renting an A100-class box beats list API prices.
From inputs to a decision
Estimate when self-hosted GPU hours beat API list prices for a given throughput. Use it before committing to A100/H100 capacity for vLLM.
- 01
Enter GPU hourly cost and utilization.
- 02
Set tokens/sec sustained throughput.
- 03
Pick API models to compare break-even against.
- 04
Read months-to-break-even and monthly delta.
Tips for better results
- Include idle capacity — utilization below ~40% stretches break-even hard.
- Use the same prompt pack and RPS when comparing models.
- Estimates use list prices — negotiate enterprise rates separately.
Frequently asked questions
- Is GPU TCO only about chip rental?
- This tool focuses on GPU $/hr vs API token spend. Add networking, ops, and egress in your own spreadsheet for full TCO.
Related tools
Keep measuring in the same cluster — or jump to the next decision.
- RAG Cost EstimatorProject monthly embedding + retrieval + generation spend for a RAG stack.Open →
- LLM Cost CalculatorSide-by-side monthly token cost across DeepSeek, OpenAI, Anthropic, Gemini, and vLLM.Open →
- Model Router RecommenderMap task types to primary/alternative models with blended $/mo estimates.Open →
- Prompt Token DiffCompare two prompts on tokens and monthly $ across models.Open →