← Tools

B2B / Engineering

Free tool

GPU / vLLM TCO

Estimate when self-hosted GPU hours beat API list prices for a given throughput. Use it before committing to A100/H100 capacity for vLLM.

  1. 1Fill inputs
  2. 2Run
  3. 3Copy / export

Inputs

Results

GPU vs API break-even

See whether renting an A100-class box beats list API prices.

How it works

From inputs to a decision

Estimate when self-hosted GPU hours beat API list prices for a given throughput. Use it before committing to A100/H100 capacity for vLLM.

  1. 01

    Enter GPU hourly cost and utilization.

  2. 02

    Set tokens/sec sustained throughput.

  3. 03

    Pick API models to compare break-even against.

  4. 04

    Read months-to-break-even and monthly delta.

Tips for better results

  • Include idle capacity — utilization below ~40% stretches break-even hard.
  • Use the same prompt pack and RPS when comparing models.
  • Estimates use list prices — negotiate enterprise rates separately.

Frequently asked questions

Is GPU TCO only about chip rental?
This tool focuses on GPU $/hr vs API token spend. Add networking, ops, and egress in your own spreadsheet for full TCO.

Keep measuring in the same cluster — or jump to the next decision.