Chapter 2
The model has amnesia between calls
When you type a follow-up question into an AI chat window, the machine answers as though it remembers everything you discussed.
If you ask for the fourth-quarter revenue of Acme Industrial Technologies, the screen prints $4,280 million. When you immediately ask, What drove that growth?, the model answers with details on equipment pricing and customer demand. It does not ask which company you mean or which earnings report you are discussing. To anyone watching the screen, the software appears to hold an active, ongoing discussion in its memory.
It does not.
To make sense of what happens on your monthly bill and on your screen, you need to understand one simple physical fact: the language model and your conversation history are two completely separate things. The model possesses zero memory. The ongoing conversation you experience is an illusion manufactured entirely by the harness.
The conversation history lives in your tab, not in the model
When people talk about a language model, they imagine a colleague sitting across a desk, keeping track of an ongoing discussion.
If you ask five questions about Acme Industrial Technologies, the tool answers with total familiarity. But if you open a new browser tab and ask about the exact same company, the tool greets you as a total stranger. People often assume the software had a temporary glitch or lost its memory.
The truth is simpler: the memory never existed on the other end.
The model retains nothing from past calls. Every time you send it a message, it starts from a completely blank slate.
How, then, does the chat window appear to remember anything?
The harness maintains the conversation history, not the model. The harness on your laptop keeps a running notebook of the discussion. When you type a follow-up question, the harness does not send your new sentence by itself. Instead, it opens its notebook, copies the entire transcript from the very beginning, pastes your new question at the bottom, and sends the whole bundle across to the model.
When the model receives this bundle, it reads the whole story from top to bottom. It sees your first question, sees the answer it gave earlier, and sees your new follow-up. It uses that entire page to predict the next words, returns the answer, and promptly forgets everything again.
When you close your browser tab, you throw the notebook away. Nothing was ever saved on the other side.
Short follow-up questions are the most expensive messages you send
Every intuition from daily business life says short requests are cheap. If you ask a junior analyst a one-sentence question, it takes them ten seconds. In an AI chat thread, the opposite happens: a fifteen-word follow-up question is often the most expensive message you send in the entire discussion.
In Chapter 1, we saw that providers bill strictly by the number of tokens processed. In an isolated, one-time question, you pay once for the words you send and once for the words you receive. The cost is small and predictable.
In an ongoing conversation, however, you pay for every past word over and over again on every turn. With each message you send, the harness re-transmits the complete conversation history alongside your newest question.
Imagine an analyst reviewing Acme Industrial Technologies across ten questions:
On the first question, the analyst sends a 1,000-token earnings excerpt and receives a 200-token summary. On the second question, the analyst types a short 50-token follow-up. The harness does not send 50 tokens. It re-sends the original 1,000 tokens, the 200-token answer, and the new 50-token question. You are billed for 1,250 tokens on the second click.
By the tenth question, the harness is re-sending all nine previous questions and answers before the model can generate a single word:
Question 1: [Excerpt] --> 1,200 tokens billed
Question 2: [Excerpt + Answer 1 + Question 2] --> 1,450 tokens billed
Question 3: [Turn 1 + Turn 2 + Question 3] --> 1,700 tokens billed
...
Question 10: [Cumulative Transcript of Questions 1 to 9] --> 3,500 tokens billed
Instead of paying for that opening financial document once, your firm pays for it ten times. If an analyst pastes a 50-page credit agreement into an early prompt, that entire filing rides along with every single follow-up question, draining your budget and slowing down response times.
In practice, some modern cloud providers offer discounts for cached text that repeats from previous messages, and some chat apps trim older turns. But the mechanical direction of the bill never changes: every new question must carry the weight of the discussion that came before it.
Running scenario: Follow up on the earnings release
In Chapter 1, you began an evaluation of Acme Industrial Technologies by prompting the model to summarize its fourth-quarter earnings release. The model performed cleanly: it reported fourth-quarter revenue of $4,280 million, noted a 6% year-over-year growth rate, and recorded operating income of $640 million.
Now put yourself in the seat of an analyst building out a research note. You do not close your browser tab and start over from scratch. You keep the chat window open and type a natural follow-up question directly below the summary:
What was the operating margin for the fourth quarter, and what drove operating income expansion?
To your eyes, you sent an eleven-word follow-up. But behind the screen, your chat harness does something much larger. It copies the entire previous exchange, packages it into a single document, and submits the full transcript to the model:
Document:
"Turn 1 Document:
'Acme Industrial Technologies (NYSE: AIT) delivered a resilient finish to the fiscal year despite persistent supply chain friction in European markets. For the fourth quarter ended December 31, consolidated revenue reached $4,280 million, a 6% increase compared to $4,038 million in the prior-year period. Operating income expanded to $640 million from $595 million, reflecting price realization across heavy equipment lines. Full-year backlogs closed at record levels, though management noted cautious capital expenditure trends among regional industrial customers entering the first quarter.'
Turn 1 Instructions:
Summarize the fourth-quarter performance for Acme Industrial Technologies. State the total revenue and the year-over-year growth rate. Use two concise sentences.
Turn 1 Output:
Acme Industrial Technologies reported fourth-quarter revenue of $4,280 million, representing a 6% increase year-over-year. Operating income reached $640 million, supported by price realization across heavy equipment lines."
Instructions:
What was the operating margin for the fourth quarter, and what drove operating income expansion?
Use two concise sentences.
The model reads the replayed transcript and generates its response:
Fourth-quarter operating margin stood at approximately 15%, calculated from $640 million in operating income on $4,280 million in consolidated revenue. Operating income expansion was driven by price realization across heavy equipment lines despite regional market friction.
The specimen number, $4,280 million, arrived intact for a second time. The model calculated an operating margin of roughly 15% ($640 million divided by $4,280 million) without asking you to re-upload the earnings release or remind it which corporation you were analyzing.
Because the output looks seamless, analysts assume the model remembered the earlier conversation. In reality, the machine only produced that answer because the harness quietly re-shipped the entire prior exchange in the background.
Look at what that calculation actually cost:
To get an answer to an eleven-word question about margins, the harness re-transmitted over two hundred words of earlier discussion. You paid over ninety percent of that transaction's bill to re-read text you already had on your screen thirty seconds ago.
Now multiply that across twenty questions on a fifty-page credit agreement. By the end of the afternoon, you are not paying primarily for financial analysis. You are paying an invisible, compounding tax to re-read your own chat thread.
What happens when your financial filing exceeds the model's memory buffer?
Here is the problem to think about before you move to the next chapter: what happens when you upload two complete 10-K annual reports or a 200-page debt agreement? If every message re-sends the cumulative history, does the model have infinite capacity to receive text, or is there a hard ceiling where the machine simply runs out of room?
The Leader's View: Budget for compounding token costs in multi-turn workflows
When rolling out conversational tools across a business or finance team, leaders must manage how staff use chat windows.
A chat interface encourages analysts to treat the screen like an infinite notepad. An analyst uploads a thick loan agreement, asks twenty questions, and leaves the tab open all week. In software with fixed monthly seat licenses, this habit carries no marginal cost. In tools billed per token, long chat threads quietly inflate operating invoices.
Beyond cost, cumulative conversations introduce an operational hazard: context contamination.
When an analyst holds an extended dialogue, early mistakes remain in the transcript. If the model makes an arithmetic mistake or misnames a subsidiary on question three, that error sits permanently in the notebook. On questions five and ten, the model re-reads that error and cites it as verified truth.
Operating rules to mandate
- Reset the thread between tasks: If your team uses automated scripts to process filings in batches, ensure the tool starts with a completely blank slate for each document. Never allow automated processes to daisy-chain multiple filings into one continuous conversation.
- Carry forward numbers, not conversation transcripts: When a workflow requires multiple steps, extract the specific figures into a clean spreadsheet or template first. Feed only those verified figures into the next step, rather than re-sending pages of past conversational back-and-forth.
- Start fresh sessions for distinct research tasks: Instruct analysts to open a new chat session for each distinct financial question rather than working in one endless thread. Opening a clean window clears out past errors and eliminates redundant token charges.
Questions to put to a vendor
- How does your platform handle transcript growth during extended conversations? Does your software automatically prune older history, or does it re-bill the entire conversation on every message?
- When an analyst asks a follow-up question, how does your system stop an early factual error from contaminating later answers?
Zoom-out: Why specialized computer chips cannot afford to remember you
Executives and finance leaders often ask a natural question: why can't the vendor simply save our company's numbers and customer history inside the brain of the machine?
The answer is not that the software vendor forgot to build the feature. The answer is the physical economics of high-speed silicon and electrical power.
When you query an ordinary company database, the software checks an index saved on a hard drive. The lookup consumes fractions of a milliwatt of power, finishes in microseconds, and the machine goes quiet.
A language model operates entirely differently. A single frontier model can run on a single rack of specialized server chips. But because hundreds of millions of people and thousands of enterprises send queries simultaneously, providers must scale that setup across thousands of servers in massive warehouse facilities.
At enterprise scale, that aggregate workload requires tremendous energy and capital. Storing personal conversation histories inside high-speed chip memory for millions of concurrent users would tie up valuable hardware and demand unsustainable infrastructure.
Instead, providers design the hardware for maximum stateless throughput. The chips process your incoming text, return the predicted tokens, clear the working memory, and immediately reassign the silicon to the next customer in line. The economics of scaling specialized hardware set the permanent price floor for every token. Knowing how that infrastructure scales helps leaders treat generative queries not as casual chitchat, but as a deliberate operating expense.