From Prompt to Bill: The Life of One LLM Call
One line of code, half a second of waiting, and a bill at the end of the month. What actually happens in between, from tokens to tail latency.
A developer writes one line of code.
client.chat.completions.create(
model="...",
messages=[{"role": "user", "content": "..."}],
)Half a second later, words start appearing. A few seconds later, the answer is complete. At the end of the month, a bill arrives.
Between that line and that bill, the request is taken apart, read, written, gambled on, forced into shape, timed and priced. Every behaviour that looks strange from the outside happens somewhere along that path: models that can’t count letters, a first word that’s slow to arrive, a temperature=0 that still changes its answer, a “1-second model” inside a feature that takes five.
This is the story of that one request.
The model can’t read
The model never sees text. It sees numbers.
"Hello world" → [13225, 2375]Two integers. The model has never seen the letter H. It has seen the number 13225 a few billion times.
Each of those numbers is a token: one entry in a fixed list called the vocabulary. Modern vocabularies hold around 200,000 entries. The tokenizer chops incoming text into pieces from that list and replaces each piece with its ID. Only the IDs go to the model.
So where did the list come from? Nobody wrote it. An algorithm called byte-pair encoding (BPE) built it from a huge pile of text. It starts with single bytes. It finds the pair of neighbours that appears most often, merges them into a new entry, and repeats. t + h becomes th. th + e becomes the. A space plus the becomes the. After 200,000 merges, the vocabulary is done.
Common words end up as one token. Rare words stay in pieces.
"strawberry" → ["st", "raw", "berry"]That one line explains a famous failure. Asked “how many r’s are in strawberry?”, the model gets three numbers and no letters. It has to remember the spelling of token berry. It isn’t looking. It’s recalling.
Now the request is a list of numbers. Time to read it.
Reading everything at once
The model takes the full list of token IDs and processes all of them together. This first phase is called prefill.
Its main job is one problem. Consider:
the animal did not cross the street because it was too tired
the animal did not cross the street because it was too wideIn the first sentence, “it” is the animal. In the second, it’s the street. Same word, same position. The only difference is a word several positions later.
So each token needs a way to pull in information from the other tokens around it. That mechanism is attention.
Run the first sentence through a real model and look at where it puts its focus:
tired 0.090 ##########
animal 0.078 #########
because 0.076 #########
street 0.049 #####Every token gets a weight, and the weights sum to 1. The model builds a new version of it by mixing the other tokens in those proportions. That’s attention: a weighted average over every other token, where the model learned how to choose the weights.
Query, key and value
How does it decide that animal deserves more weight than street? Every token produces three vectors, and each one answers a different question.
The query answers “what am I looking for?” For it, the query roughly says: I’m a pronoun. I need the noun I refer to, something that can be tired.
The key answers “what do I offer?” Every other token puts one up.
animal → "I'm a noun. A living thing."
street → "I'm a noun. A place."
tired → "I'm a state a living thing can be in."The model compares the query of it against every key. A close match gives a high score, so animal scores higher than street. Those scores become the weights you saw above.
The value answers “if you pick me, what do I hand over?” For animal, that’s its meaning: a creature, the subject of the sentence. The weights decide how much of each value flows into it. After the mix, it carries information about the animal.
query of "it" · key of each token
→ scores
→ weights
→ weighted mix of valuesQuery and key decide how much. Value decides what.
The sentences in quotes are human translations. In the model, each query, key and value is a list of learned numbers, not words. A model also runs many of these comparisons side by side, called attention heads, each learning to look for something different: one for pronouns, one for verbs, one for position.
The catch
Every token compares itself with every other token. So the work grows with the square of the prompt length.
100 tokens → 10,000 pairs
1,000 tokens → 1,000,000 pairs
10,000 tokens → 100,000,000 pairsDouble the prompt, quadruple the attention work. This one fact is why context windows stayed small for years, and why retrieval (RAG) exists at all.
Prefill is compute-bound. Thousands of tokens are processed in parallel, and the GPU’s maths units are the limit. All of it has to finish before the first output word can exist. So prefill decides how long a user waits before anything appears. That wait is called time to first token (TTFT), and it comes back later when we talk about latency.
The prompt has been read. Now the model has to say something.
One word at a time
Once prefill is done, the model moves to its second phase: decode. This is where the answer gets written.
It produces the first token, adds it to the input, then runs again to produce the next token. This continues until the response is complete. Writing one token at a time, each one fed back in as input, is called autoregressive generation.
Running the full attention calculation again for every earlier token, each time, would be wasteful. The earlier tokens haven’t changed, so their keys and values are still the same. The model keeps those keys and values in memory, in what’s called the KV cache. When a new token arrives, the model only calculates the new token’s attention against the cached keys and values. Nothing is recalculated from scratch.
This makes decode cheap in maths. It also makes it slow in a different way.
To produce a single token, the GPU has to read every weight of the model from memory. For a 27-billion-parameter model at 16-bit precision, that’s roughly 54 GB of reading for one token’s worth of maths. The maths finishes almost instantly. The reading is the bottleneck. Decode is memory-bandwidth-bound.
And because tokens come out one by one, decode time grows with how much the model writes.
That leaves every request with two separate waits.
The first is the wait for the first word. That’s prefill, so it grows with the length of the prompt.
The second is the wait for the last word. That’s decode, so it grows with the length of the answer.
So if the first word is slow to appear, shorten the prompt. If the answer takes too long to finish, shorten the output. Same request, two different problems, two different fixes.
But “produce one token” hides a decision. Which token?
A roll of the dice
Decode produces one token at a time. But the model never writes a token directly.
Remember the vocabulary, the list of ~200,000 tokens? At every step, the model gives every one of them a score. Those raw scores are called logits.
sunny 4.2
nice 3.9
cold 3.1
purple -1.2
...and ~200,000 moreLogits aren’t probabilities. They can be negative, and they don’t add up to anything. A function called softmax turns them into probabilities that add up to exactly 1:
sunny 41.2%
nice 30.5%
cold 13.7%Then one token is picked from that distribution. This whole process of picking one token out of the entire vocabulary is called sampling.
What temperature really does
People often describe temperature as a “creativity” setting. It isn’t. The model has no creativity dial. Temperature changes how the probabilities are spread out before the pick, and the whole mechanism is one division:
probs = softmax(logits / temperature)temperature 0.1 → sunny 95.3% nice 4.7%
temperature 1.0 → sunny 41.2% nice 30.5%
temperature 2.0 → sunny 28.7% nice 24.7%With a low temperature, the token sitting at the top with the highest probability almost always gets picked. With a high temperature, everything flattens toward equal, and more options get a fair chance of being picked.
So “creativity” isn’t something the model has. It comes from how much randomness we allow when choosing the next token.
Cutting the list: top-k and top-p
Even after temperature, the distribution still covers all ~200,000 tokens. Most have tiny probabilities, but tiny isn’t zero. Generate enough tokens and one of those junk tokens eventually gets picked. So before the pick, the candidate list is usually trimmed.
top-k keeps only the k most likely tokens and throws away the rest. With top_k = 3:
sunny 48.2% nice 35.7% cold 16.1%
← the 3 survivors, rescaled to add up to 100%top-p, also called nucleus sampling, keeps the smallest group of tokens whose probabilities add up to p. With top_p = 0.9:
sunny 41.2% running total 41.2%
nice 30.5% 71.7%
cold 13.7% 85.4%
warm 10.1% 95.5% ← crossed 90%
stop here: 4 tokens keptThe difference shows up when the model’s confidence changes. When the model is 99% sure, top-p keeps one token, but top-k still lets in k. When the model is torn between many options, top-p keeps many, but top-k still caps at k. top-p adapts to the model’s confidence; top-k can’t. That’s why top-p is the more common choice.
Temperature 0 isn’t deterministic
At temperature=0, providers skip sampling entirely and take the single highest logit. This is greedy decoding. No randomness is involved.
And yet the same prompt still returns different answers. Here’s why:
sunny 4.20013
nice 4.20011 → greedy picks "sunny"
sunny 4.20009 ← nudged by 0.00004 of floating-point noise
nice 4.20011 → greedy picks "nice"Sampling didn’t change. The inputs to the pick wobbled. GPUs add numbers in different orders depending on how a request is batched with other users’ requests. Floating-point addition isn’t associative, so the order changes the last digits. Mixture-of-experts routing and silent model updates add more. Because generation is autoregressive, one flipped token changes everything after it.
Every token is a probability draw. So how does anyone get reliable output from this?
Getting the response according to our requirements
Imagine your code can only accept JSON. Not prose, not “Sure! Here’s your JSON:”, just a JSON object that parses every time. How do you make a probability machine produce that?
There are four ways, from weakest to strongest.
1. Ask nicely. Put “respond in JSON” in the prompt. The model usually complies, and that’s the problem: usually.
2. JSON mode. Providers now offer a JSON mode you switch on in the request. It guarantees the response is valid JSON. It does not guarantee it’s the JSON you wanted. Keys can be missing, renamed, or the wrong type.
3. Structured output. This one guarantees the exact shape. You give the provider a JSON Schema, a description of exactly which keys you want and what type each one is. In Python it’s usually generated from a Pydantic class:
class Invoice(BaseModel):
vendor: str
amount: float
response_format = {
"type": "json_schema",
"json_schema": {
"name": "invoice",
"schema": Invoice.model_json_schema(),
"strict": True,
},
}This is where the dice come back. The provider turns the schema into rules about what’s allowed to come next at every position. Then, during decode, right before each pick, it checks every candidate token against those rules. Any token that would break the schema has its probability set to exactly zero.
output so far: {"vendor": "Acme", "amount":
candidate allowed?
" 42" yes a number can start here
" \"" no amount must be a number, not a string
" hello" no not a number
"}" no amount needs a value firstThe model can’t produce invalid output, because invalid tokens are never on the menu. This is called constrained decoding.
4. Tool calling as extraction. Define a “tool” whose arguments are the fields you want. The model returns a tool call, which is a request, not an execution. Nothing runs. The arguments are your data.
Even a perfect schema can’t check business rules: totals that must add up, dates that must be in the past. Those need a validator in code, and a retry that sends back both the invalid output and the error message.
The output is now shaped and correct. But a user was waiting the whole time. How long did it take?
The clock
The numbers teams actually watch
In production, LLM latency is measured with two numbers.
Time to first token (TTFT) is how long the user waits before the first word appears. That’s prefill, plus any waiting in line before it.
Total time, also called end-to-end latency, is how long until the last word arrives. That’s TTFT plus all of decode.
For a chat feature, TTFT is what users feel. That’s why the single biggest latency lever doesn’t make anything faster. Streaming shows tokens as they arrive. A 10-second answer that starts in half a second feels fast. The same answer shown all at once feels broken.
For a background job, nobody is watching. Latency barely matters there. What matters is throughput, how many jobs finish per hour, and cost.
Latency is a distribution, not a number
Send the same request 20 times and you get 20 different times. Here’s an example run, sorted from fastest to slowest:
#1 0.61s #6 0.68s #11 0.75s #16 0.88s
#2 0.63s #7 0.70s #12 0.77s #17 0.93s
#3 0.64s #8 0.71s #13 0.79s #18 1.02s
#4 0.66s #9 0.72s #14 0.81s #19 1.40s
#5 0.67s #10 0.74s #15 0.84s #20 3.10sA percentile is read straight off this sorted list. For the p-th percentile, take p% of the number of samples and read the value at that position.
p50 → 50% of 20 = position 10 → 0.74s
p95 → 95% of 20 = position 19 → 1.40s
p99 → 99% of 20 = 19.8, round up to position 20 → 3.10sp50 is the median. Half the requests were faster than 0.74s. p95 means 95% were faster than 1.40s, and one in twenty was slower. The slow end of the list is called tail latency, and it’s what users remember.
Now the average: all 20 added up is 18.05s, divided by 20 is 0.90s. Sixteen of the twenty requests were faster than that. The average describes almost nobody. One slow request at 3.10s dragged it up, and it still hides how bad that one was.
Fan-out: why the tail takes over
Real features almost always make more than one call. A RAG answer might embed, retrieve, rerank and generate. An agent might call the model twenty times to finish one task.
If each individual call has only a 5% (1-in-20) chance of being slow, that sounds rare. But when one page or task makes many calls, the chance that at least one is slow becomes much higher.
The chance that one call is fast is 0.95. The chance that all n calls are fast is 0.95 multiplied by itself n times. Everything else is “at least one was slow”.
With 5 calls:
1 − 0.95⁵ ≈ 22.6%
So roughly 1 in 4 pages hits at least one slow call.
With an agent making 20 calls:
1 − 0.95²⁰ ≈ 64%
So most agent tasks, about 64%, hit at least one slow call.
That’s fan-out: one user action depends on many downstream calls, so even rare slowdowns compound. It’s why the median can mislead. Your typical call might be fast, while the slowest calls, the p95 and p99 tail latency, decide how slow the whole experience feels. That’s why teams put p95 and p99 on their dashboards, not the median.
The critical path
When a feature makes several calls, some depend on each other and some don’t. Calls that don’t depend on each other shouldn’t wait on each other.
three calls, one after another 1.90s
call 1, then calls 2 and 3 together 1.16s (39% faster)The critical path is the longest chain of steps that has to happen in order. Only shortening that chain makes the whole thing faster. Speeding up anything off that chain changes nothing the user can feel.
The latency budget
A feature isn’t one call. It’s a chain of steps, and the LLM is only one of them. So when someone says “the feature is slow”, the first question is: which step?
A latency budget answers that before it becomes a problem. The team first decides the total wait users will accept. Then they split that total across every step, and hold each step to its share. For a chat reply that should start within two seconds:
embed the question 150 ms
vector search 100 ms
rerank 250 ms
LLM time to first token 1,200 ms
network + app 300 ms
-----------------------------------
2,000 msWhen the feature gets slow, the budget shows which step went over. When someone wants to add a step, the budget shows what it has to take time from.
The LLM is one line in that table. In a real system, DNS lookups, TLS handshakes, cold starts, database queries and queueing behind other requests all claim their share too.
The response arrived, and now we know how long it took. One question is left: what did it cost?
The bill
Latency is what users feel. Cost is what the company feels.
LLM providers charge per token, with separate prices for input tokens (what you send) and output tokens (what the model writes). Prices are quoted per million tokens, and output is usually several times more expensive than input.
Before building a feature, teams estimate what it will cost to run each month. The estimate is built from five things: how many requests the feature gets, how many LLM calls each request makes, how often calls fail and get retried, how many tokens go in and out per call, and the price of those tokens.
monthly LLM cost of a feature
= requests per day
× LLM calls per request
× (1 + retry rate)
× 30 days
× (input tokens × input price + output tokens × output price)Two inputs get forgotten most: calls per request (it’s rarely one) and the retry rate (every retry is paid for again).
An example: a support-reply feature. It gets 20,000 requests a day, makes two calls each, sends 2,500 input tokens, gets 400 output tokens, retries 5% of calls, and uses a model priced at $2 per million input tokens and $10 per million output tokens.
calls per month 20,000 × 2 × 1.05 × 30 = 1,260,000
cost per call 2,500 × $2/1M + 400 × $10/1M = $0.009
monthly cost 1,260,000 × $0.009 = $11,340Most of those 2,500 input tokens are the same in every request: the system prompt, the instructions, the examples. Prompt caching stores the prefill work for that fixed beginning of the prompt, and cached input is billed at about a tenth of the price. If 80% of the input is cacheable, the bill drops to $6,804, which is 40% less.
That’s the rare lever that helps both the clock and the bill. Skipping prefill makes the first token arrive sooner, and paying a tenth makes the bill smaller. It costs nothing in quality.
The whole journey
tokenizer text → token IDs
prefill attention over the whole prompt sets TTFT
decode one token at a time, KV cache sets total time
sampling logits → softmax → temperature → top-p → pick
constraint schema zeroes out invalid tokens
the bill tokens × price × calls × retries, felt at p95Every strange behaviour on the outside has a home somewhere on the inside. The model can’t count letters because it never saw letters, only numbers. Long prompts are slow to start because every token had to look at every other token first. Long answers are slow to finish because the words come out one at a time. temperature=0 wobbles because the dice were never the only source of chance. A schema can’t stop a made-up PO number, but an optional field can, because a required field is an order to produce something. And a fast model makes a slow feature because users don’t live at the median.
Knowing where along the path a problem lives is most of the work of fixing it.