In my 2026 AI models showdown I treated giant context windows as a spec to compare. This time I want to look at what they do to real work, and AI in journalism is the best test case I know. Reporters, fact-checkers, and researchers share one problem: too much to read, and a deadline that doesn’t care.
My take up front: a million-token model makes reading faster. It doesn’t make anything true. Every good example I found keeps a human checking the output.
What a Context Window Is (and a Correction)
A context window is how much text a model can hold in view at once. It’s measured in tokens, which are chunks of text a bit smaller than a word.
Now the correction. In the showdown I listed Gemini 3.1 Pro at 2 million tokens. Google’s own model documentation lists 1,048,576. OpenAI’s lists 1,050,000 for GPT-6 Astra, and Anthropic’s lists 1 million for Claude Opus 5.5. So the three flagships are level at roughly a million tokens, and my earlier comparison was wrong on that point.
A million tokens is roughly 750,000 English words. Picture well over a thousand pages of meeting minutes in a single prompt.
One caveat. Fitting a document in the window doesn’t mean the model reads every page with equal care. A 2025 study by Chroma tested 18 models and found they all got less reliable as the input got longer. Those were last year’s models, and Chroma sells search tools, so take it with some salt. I still haven’t seen anyone show the problem is gone.
AI in Journalism: Sorting the Pile Faster
The clearest evidence comes from this year’s Pulitzer Prizes. Nieman Lab reported that five winners and three finalists disclosed AI use to the judges, the most since the Pulitzers started asking in 2024. Most used large language models (LLMs, the technology behind chatbots) for the same job: getting through piles of documents.
| Newsroom | What the AI did | What humans still did |
|---|---|---|
| The Wall Street Journal | Summarized every page of county meeting records after the 2025 Texas floods | Read every flagged section, then every relevant document in full |
| The Minnesota Star Tribune | First-pass translation of a coded journal, more than 600,000 words | Sent key passages to two language academics, who caught errors |
| Associated Press | Made tens of thousands of leaked emails and records searchable | Reviewed the documents by hand and never quoted the AI summaries |
| The New York Times | Double-checked how reporters had classified more than 10,000 documents | Read and classified everything manually first |

Look at the right-hand column. Nobody let the model decide what was true. John West, a computational journalist at the Journal, told Nieman Lab the goal was to “sort the pile so the most relevant stuff is right at the top.”
There’s a catch for readers. According to the same report, the Journal’s flood stories carried no AI label. The paper’s reasoning was that the tool worked like a search engine and reporters read the documents themselves. Fair enough, but it means you often can’t tell from the article.
What to do with this
- If you read the news: look for a methodology note on big investigations. Outlets that explain how they used AI are showing their work.
- If you do the work: let the model rank and summarize, then read the originals of anything you plan to publish.
Source Verification: The Model Can’t Vouch for Anything
This is the part where I’m most skeptical.
The Reuters Institute’s Digital News Report 2026 found weekly use of AI chatbots for news rose from 7% to 10% in a year. Among those users, 33% ask the chatbot to assess whether a news source is reliable. Only 20% of the general public say they trust news from AI chatbots.
So a third of AI news users treat the chatbot as the referee, and its own record on news is shaky. An October 2025 study led by the BBC and the European Broadcasting Union had journalists grade more than 3,000 AI answers about the news. They found at least one significant issue in 45% of them. Those were 2025 assistants and I couldn’t find a newer edition, so read it as a warning, not a current score.
A bigger context window doesn’t fix this. A model can compare ten documents and tell you they agree. It can’t tell you whether the leaked email is forged or the witness lied. Verification means checking a claim against the world, and the model only has the text you gave it.
The Star Tribune case shows why. The AI translation was mostly right, but the human experts found a passage that misrepresented the attacker’s possible motive. That’s exactly the kind of line that ends up in a headline.
The Times did it the other way round, and I like that better. Humans reviewed everything first, and the model gave a second opinion. Where the two disagreed, reporters went back and read the documents again.
What to do with this
- If you read the news: when a chatbot summarizes a story, open the original before you repeat it.
- If you do the work: ask for the exact passage and page behind every claim, then go and look at that page.
- Both: never let the same model write the claim and grade the claim.
Pro Tip: Give the model two sources that should agree and ask it to list only where they differ. A contradiction is quicker to check than a summary, and it’s usually where the story is.
Academic Research: Fast Reviews, Fake Citations

For students and researchers the appeal is obvious. Load fifty papers into one prompt and ask what they agree on. That used to be a week of reading.
The damage is already measurable. An audit published in The Lancet checked about 2.5 million biomedical papers and found 4,046 fabricated references spread across 2,810 of them. In 2023, about one paper in 2,828 had at least one made-up reference. In the first seven weeks of 2026, it was one in 277.
The authors say the jump lines up with the spread of AI writing tools. They don’t claim every fake came from a chatbot. Paper mills, which are businesses that sell fake studies, are in the mix too, and the audit only covered open-access papers.
This is where a big context window can help, if you use it the right way round. A model asked to produce references from memory is guessing. A model asked about PDFs you loaded yourself is at least reading something real. It can still misread them, so the checking stays.
What to do with this
- If you read research news: “a study found” deserves a link. If there’s no link, don’t pass it on.
- If you do the work: load the sources yourself. Don’t ask the model to find them from memory.
- If you do the work: open every reference before it goes into your bibliography.
The Verdict
Big context windows are the most useful thing to happen to reading-heavy work in years. They’ve changed nothing about who is responsible for the result.
The newsrooms the Pulitzers recognized this year use AI to find the page worth reading. Then a person reads the page. If a tool saves you a week of reading, spend an hour of that week checking what it told you.
Frequently Asked Questions
What is a context window in simple terms?
It’s the amount of text an AI model can consider at one time. The current flagship models list about a million tokens, which is roughly 750,000 English words.
Is AI in journalism replacing reporters?
Not in the work that was recognized by the Pulitzers this year. Those newsrooms used AI to sort and summarize documents, and reporters still read the sources and verified the findings.
Can AI fact-check the news for me?
It can point you to passages and spot contradictions between documents. It can’t confirm that a document is real or that a source told the truth, so you still have to check the original.
Why do AI tools invent citations?
A language model predicts plausible text. Asked for a reference it doesn’t have, it can produce one that looks right and doesn’t exist. Giving it the real documents cuts down the guessing.
Do newsrooms have to tell readers when they use AI?
It depends on the outlet and how the AI was used. The Pulitzers require entrants to disclose it to the judges, but some outlets don’t label stories when AI only helped search documents.


