Prickled.ai ← Back to Prickled.ai
ROMTI · Return on Model Token Investment

You are paying by the word. Nobody told you the rate.

Paste a real prompt. Pick your model. See what that one request costs, what a full conversation costs, and what a month of the same habit costs across Claude, Gemini, Copilot, ChatGPT and Grok.

Nothing you type leaves your browser. No prompt is sent anywhere. All math runs on this page.

ROMTI is Return on Model Token Investment. Not what the model costs. What the model returns for what it costs, once you count the tokens you never see and the hour you spent steering it. ROMTI = (value of the output − your time cost) ÷ total token cost

The essay behind this page: Return on Token Investment, on what a month of Fable 5 actually taught me about the bill.

A woman working at a laptop, smiling
// what an hour of this actually costs
The simulator · 01

Price out your own prompt

Type or paste what you would actually ask. The estimator counts tokens the way the models do, then prices the whole job, not just the sentence you typed.

1. Your prompt

Paste the real thing. Length is what costs money.

0 tokens in your prompt
0 characters
Try one

2. Your model

Published API rates per million tokens, checked 9 August 2026.

 
6
Every turn re-sends everything said before it. This is where the bill actually comes from.
The costs you cannot see
Charged as input on every single turn, before you type a word.
0
Roughly 750 tokens a page. A 40 page PDF is about 30,000.
0%
Cache reads bill at roughly a tenth of the input rate. Most people never turn this on.
25%
The runs you threw away. They still billed.

3. Your habit and your hour

Cost per prompt is trivia. Cost per month is a decision.

Be honest. If it replaced nothing, the ROMTI is negative no matter how cheap the tokens were.
What this conversation costs
$0.00
6 turns, including retries
One turn
$0.00
first message only
Per month
$0.00
at this habit
Tokens in
0
billed as input
Tokens out
0
billed as output
Where the money actually goes
Your prompt
$0
Re-sent history
$0
System & tools
$0
Attached files
$0
The answer
$0
Thinking tokens
$0
Retries
$0
Your actual words are 0% of the bill.
ROMTI on this conversation
Value created
$0
work replaced
Your time in
$0
steering it
Token cost
$0
what you priced
ROMTI
return per dollar
 
Same job, five vendors · 02

The identical conversation, priced five ways

Your exact prompt, your exact settings, run on a comparable model from each of the five. Choose the tier you would realistically use.

Tier

And the part the comparison hides

Most people never touch an API key. They pay a flat monthly fee. So the real question is whether your usage is worth more or less than the subscription.

 

The costs nobody prices in · 03

Three things on your bill that are not your prompt

Token pricing looks simple because vendors publish two numbers. The bill is not made of two numbers. These are the three that move it most.

Cost one

The context tax, and it is quadratic

A model has no memory. Every turn, the entire conversation so far is sent again as fresh input. Turn 10 is not ten times turn 1. It is closer to fifty-five times, because turn 10 carries turns 1 through 9 on its back.

total input over N turns = N × overhead + N × prompt + (prompt + answer) × N(N−1)/2
A 30 turn thread sends the early messages 30 times. Same words. Thirty invoices.
Cost two

The prompt you did not write

Before your first character, the app has already sent a system prompt, a persona, safety instructions, and a schema for every tool and connector it can reach. A loaded agent can carry 15,000 to 20,000 tokens of that. It is charged as input. On every turn.

Connecting nine integrations you never use is not free. It is a standing charge.
Cost three, and the largest one

Your hour costs more than every token you will ever buy

At $65 an hour, six minutes of your attention is $6.50. A full Claude Opus 5 conversation at the settings most people run costs a few cents. The token bill is a rounding error against the human bill. Which means optimising the wrong number is the actual waste.

This is why ROMTI is a ratio and not a price. A model that costs four times more and gets it right the first time is not expensive. It is the cheap option, because the expensive input was never the tokens.

And the externality: Google measured a median Gemini text prompt at 0.24 watt hours, 0.03 grams of CO2e and 0.26 millilitres of water. Small per prompt. Not small at eight a day, forever, times everyone.
The stance
Cheap tokens spent badly are the expensive option.

Nobody wins this by picking the lowest price per million. They win it by needing fewer turns, sending less dead weight, and knowing which jobs deserve a frontier model and which deserve the cheap one. That is a skill, and it is learnable in an afternoon.

The rest of it

The part that actually changes your bill

The simulator above is free and always will be. What is below it is the analysis: where the money really goes, and the changes that move it.

01
Seven costs on your bill that are invisible in the price list. Re-sent history, system prompts you never wrote, reasoning tokens, retries.
02
Ten changes, ordered by what each one saves you. All free to do today. The savings compound.
03
The nine ideas behind every number on this page. Enough to argue with a vendor and win.
04
Your scenario, written up as a report you can hand to someone. Priced across all five vendors, with the three changes that would cut your bill most.
One email with your report. No sequence, no list swap. Your prompt text is never included or transmitted.

A bill is a symptom.

What it usually means is that nobody ever designed the stack underneath it, so every job runs on whatever tool happened to be open. That is fixable in one sitting. Your scenario is where it starts: I audit what you actually run on, then build the stack around the work you actually do.

Bring me your numbers

Your settings become the first page of the audit. Your prompt text is never part of it, only the choices you made.