Archives: Newsletters

I onboarded a team of journalists to Claude Cowork. Here's how it went.

My husband finds my tech-adjacent title – professor of practice of AI and investigative journalism – hilarious. At work, I cheerfully wrestle with every new AI tool that crosses my path. At home, I refuse to learn the TV remote and hunt him down every time I want to switch between the news and Netflix. The technical term, I believe, is willful ignorance.

But my domestic refusal to embrace a growth mindset has given me real empathy for the complicated feelings people bring to learning AI. I know that if my husband wrote me instructions for the remote, I’d learn it in a minute. But the second he starts lecturing, I shut down completely.

My colleagues, hard at work.

So, when my boss, Mark Greenblatt, asked me to onboard my journalist colleagues onto the AI agent Claude Cowork, I brought this perspective to the task.

For context, I am obsessed with Claude Cowork. I use it daily and it has meaningfully improved my life. I was so excited for this assignment, in the way you want everyone around you to read a book you love.

I brought seven of my colleagues together for an hour-long workshop. A few weeks later, I interviewed five of them about their experiences and whether they were still using the tool. The discussions felt honest, shedding light on anxieties and breakthroughs that I suspect are broadly applicable.

In what follows, I have anonymized my colleagues, with the exception of Mark.

Bad prior experiences with AI

Except for Mark, everyone I interviewed had past negative experiences that made them skeptical about either the training, or AI in general.

A colleague I’ll call Anya described a constant internal conflict. She worried about AI generating falsehoods, and the environmental impact, yet felt a professional obligation to master it. A colleague I’ll call Maeve was skeptical after a string of frustrating encounters, and Liam said he started his job at Cronkite “predisposed to not really want to use it.” Isla was triggered by the word “agent,” wondering if the tool was going to access her email or Google Drive.

The value of seeing me use AI

Several colleagues noted that their attitudes began to shift even before the official workshop, simply from watching me use the technology in real time. Liam said his “mental approach” transformed, “passively through osmosis,” while we were working together to fix a font issue on our website. He watched me take a screenshot of WordPress, feed it to Gemini, and instantly receive the troubleshooting steps we needed to solve the problem. “That was my first experience, seeing the power of AI,” Liam said.

Training day

For the training, everyone except one remote colleague was physically in the same room. I had prepared a step-by-step worksheet designed to get everyone set up with Claude Pro and then walk them through a simple web-scraping task.

Being physically present proved essential for handling the inevitable setup friction. While I tried to anticipate as many hurdles as possible, unexpected tech glitches still cropped up. For instance, Claude performs web scraping best via its Chrome extension, but Maeve hadn’t been able to open Chrome for months due to a system error. Anya successfully installed the plugin, but Claude kept trying to route her through Safari. I spent nearly the entire workshop jumping from desk to desk to troubleshoot these kinds of issues; for most of the group, simply getting set up took over half an hour.

Mark later noted that the most valuable thing I did was getting everyone “physically in a room” because there was value in struggling through those initial tech annoyances together.

The lightbulb moment

As I walked around the room, I began to see people have the lightbulb moment. I had deliberately chosen a web-scraping exercise because it is something standard chatbots typically can’t handle, making it a showcase for the power of agentic AI. It is also a tedious chore that most people would gladly automate. And while errors are possible, it is a specific use case where hallucinations are less of a risk.

By the end of the hour, everyone had successfully completed the exercise. Liam recalled being “amazed,” and Anya called it “really cool.” Mark, who had originally pushed me to test AI agents back when I was a skeptic, but hadn’t actually used Cowork himself yet, said the workshop was “an eye opener [...] about how awesome this technology is.”

Some are still using it, some aren’t

The workshop had the greatest impact on Maeve. When I asked if she was still using Cowork, she immediately rattled off multiple examples from just that morning. She is now using it to scrape tax data, and noted that Claude even tracked down state spending information while she was on a phone call. Mark also found immediate success, using Cowork to salvage a corrupted document he had spent hours editing.

For others, however, the momentum stalled. Several colleagues don’t appear to have touched Cowork since the workshop. Isla used the agent to review dozens of public records PDFs for story ideas, but she wasn’t impressed with the results and hasn’t logged back in. Even so, she valued the training; because she reports on AI, she recognizes that she needs to deeply understand the technology to cover it effectively.

"I just worry about being dead weight"

When I asked everyone how they felt during the workshop, I expected a standard level of pushback. I thought one or two colleagues would admit they were annoyed to be pulled away from their deadlines, or skeptical of my AI evangelism. Instead, I found myself saddened by Liam’s response. He said that an anxiety had surfaced during the session. The “feeling that maybe I’m a bit of a dead weight around AI […] that was an emotion that was bubbling up as the training was happening,” he told me. “This is something I want to learn,” he said, “and I’m worried that I won’t.”

Our conversation then shifted to a specific challenge where Liam had previously tried, and failed, to use Gemini. I confidently declared it a perfect use case for Cowork, and together we started a new task. Claude, however, quickly informed us it could not, in fact, complete the request. Humbled, we wrapped up the meeting, leaving me to wonder about the need to perhaps temper my AI enthusiasm.

Viewpoint diversity

Surrounding myself with colleagues who span the entire spectrum of AI opinions keeps me grounded. One colleague’s caution, for instance, inspired us to stress-test an AI tool I had confidently declared hallucination-free, which ended up exposing a flaw. Meanwhile, I have Mark providing ideas and encouragement for responsible AI experimentation, and the Scripps Howard Foundation provided funds for testing these tools.

If the workshop taught me anything, it’s that onboarding people to AI is less about evangelism and more about presence – being in the room when it works, and especially when it doesn’t.


What we're testing

The showdown: Claude’s Cowork vs. OpenAI’s Codex

While this newsletter focused on Claude’s Cowork, many journalists only have access to OpenAI’s Codex. Both are AI agents with similar capabilities relevant for journalists – they can scrape the web, sort through massive documents, and conduct research.

To see how they stack up, I put them head-to-head in a real-world journalism use case. I gave both agents the exact same prompt: pull the tax filings for the 20 largest non-profits in Arizona, calculate the percentage of revenue spent on executive compensation versus program services, generate a tip sheet, and link to all sources. For the curious: Claude ran on Claude 4.7 Opus, and OpenAI used 5.5 High. Here’s how they performed.

The results

Claude did substantially better, but frankly, both were amazing. If I had seen Codex’s performance just last year, I would have been in awe. It is only when compared directly to Claude that it feels a bit underwhelming. But both absolutely got the job done, and then some.

In minutes, both agents generated Excel files that would have taken me hours to build from scratch, extracting key values from dense tax forms and running the calculations.

They also displayed a trait I’ve seen repeatedly from agentic AI that never ceases to amaze me: intelligent disobedience.

Like a guide dog refusing a command to walk because it sees a car speeding into the intersection, both Claude and Codex recognized the implied assumption in my prompt, that a higher executive-compensation-to-programming ratio is inherently bad. Instead of just blindly running the numbers, both agents flagged cases where executive compensation was listed as zero and warned me that this wasn’t necessarily a good thing.

Intelligent disobedience in action: Both Claude’s Cowork (top) and Codex (bottom) refused to blindly follow the prompt, instead flagging how misleading a reported “$0” in executive compensation can be.

Where Claude took the win

Ultimately, Claude beat Codex across several key dimensions:

Better links to sources: My prompt explicitly asked to “link to all sources to facilitate verification.” Codex provided links, but they just pointed to a massive Excel sheet containing tens of thousands of rows of raw tax data. For true verification, a journalist wants the exact URL where the data originated. Claude understood this. It created a dedicated column in its Excel sheet linking directly to the specific tax filing on ProPublica.

Cleaner data formatting: Claude’s spreadsheet was vastly easier to eyeball and digest at a glance compared to Codex’s output.

Better editorial insight: Claude’s tip sheet was more detailed, easier to make sense of, and surfaced more compelling journalistic leads.


Westlaw’s AI features

We’ve been putting Westlaw’s AI features to the test, exploring Thomson Reuters’ version of a walled-garden chatbot. To give you an idea of how it works: Mark Greenblatt, the Howard Center’s executive editor, told me about a student who was hitting a brick wall on PACER, hunting through local jurisdictions for lawsuits related to a story. Mark tried typing the company name directly into Westlaw’s AI search, asked if there were any connected lawsuits, and multiple cases popped up in minutes. Mark told me that, “the research that we were able to do gave new life to his entire investigation, and he walked away really, really happy.”

The catch? It’s pricey . AI-assisted plans start at $155 a month per seat. But according to Mark, the time saved can make it worth the price for some use cases. It’s great at spotting trends and connecting the dots between federal and local cases fast. He also pointed out a huge perk for investigative work: if a government agency denies a records request using a specific exemption, the tool lets you easily hunt down past records-related rulings. Finding those precedents makes it much easier to build an appeal.

I decided to give it a test drive myself with a curveball question from our recent investigation into cosmic radiation protections for US aircrews. We had found one regulation requiring airlines to have “a plan for mitigating crew exposure to radiation during solar flare activity” on polar flights. Our team had gone back and forth on whether this actually counted as a real protection, so I wanted to see if Westlaw would catch the nuance and at least flag the provision. Sadly, it didn’t. Like so many AI tools, it’s not perfect, but it can be a time-saving assistant.

A screenshot from Westlaw’s AI-Assisted Research tool.

Could journalists’ AI prompts be subpoenaed? An interview with media attorney David J. Bodney

An interview with media attorney David J. Bodney

Discussions about the risks of AI often focus on the tendency of AI models to make stuff up. For journalists, however, there is an additional risk: a subject of a journalist’s investigation could ask a court to issue a subpoena to an AI company, compelling it to turn over the reporter’s prompts.

A prompt is any input provided to an AI model. It might be a question, an instruction, or even a file upload. How should reporters think about this risk?

David J. Bodney
David J. Bodney is Senior Counsel at Ballard Spahr

We interviewed David J. Bodney to find out. David is Senior Counsel at the law firm Ballard Spahr. He founded the firm’s Media and Entertainment Law Group, and has defended major news organizations in First Amendment and privacy cases for over 40 years, and briefed high-profile cases before the U.S. Supreme Court.

David has served as adjunct faculty at Arizona State University’s Walter Cronkite School of Journalism and Mass Communication, Sandra Day O’Connor College of Law and at the James E. Rogers College of Law at the University of Arizona.

When prompts become evidence

BODNEY: I think there are all kinds of risks associated with the use of AI. And the subpoena of a reporter’s prompts is really just one of those risks.

The successful subpoena of such information would be a roadmap into the reportorial process. And it should be protected under the First Amendment, depending on the circumstances in which the subpoena arises.

If it’s a third party subpoena, say, in civil litigation, the reporter might stand a better chance of protecting the work product. If the reporter is a defendant in a defamation case, and the request for the documents or information arises in that context, the likelihood of a strong First Amendment defense, I think, decreases.

Proving intent and malice

BODNEY: If you’re revealing the reporter’s intention in the prompt, an adversary could deduce from the language used that the reporter had a bias, or worse.

The prompt could demonstrate malice, common law malice, which in Arizona is defined as hatred, contempt or ill will, as opposed to constitutional actual malice, which is knowledge of the falsity of reckless disregard for the truth.

That’s particularly risky. The prompt could show bias, it could show malice, it could reveal lots of things about intention, and intention is hugely important, you know? And a contemporaneous writing may be powerful proof of a journalist’s intention, more so than a lot of hemming and hawing on a witness stand two years later.

“If you’re revealing the reporter’s intention in the prompt, an adversary could deduce from the language used that the reporter had a bias, or worse.”

The illusion of privacy

BODNEY: Only if journalists are unaware of those risks are they
uniquely risky, because they are no different really from other kinds of communications.

I’ll give you this example. In defamation litigation where I have defended news organizations, there’s nothing worse than having to review 10 journalists’ email communications over a two-year period of time to see what’s in there. And quite possibly what’s in there in the form of written communications could be far more damaging to the case than the sentences at issue in the lawsuit itself.

Claude Prompt
A hypothetical Claude prompt. David J. Bodney says prompts can be a roadmap for the reportorial process.

Why? Because the reporter has gotten familiar with their keyboard and believes their computer is their trusted friend. Say, they’re emailing a former colleague who’s now at a different news organization. And they are sharing their innermost thoughts, and they’re thinking, incorrectly, they will forever be protected.

But the reality is that it all can be discovered. Or at least that should be the guiding assumption – that everything one puts in writing as part of the newsgathering process could be discovered. It may be subject to a First Amendment or some other privilege, but it may not, particularly if you are the subject of litigation.

I think the prompt is akin to: what did you ask the source? What did you say to the source? How did you phrase the question? And I think there’s a risk of becoming, or being viewed as, negligent, you know?

Becoming too familiar with those prompts. So is that a risk? Yeah.
What’s it akin to? And that’s kind of the way the law works its logic around these issues. So, how is it different from, you know, a phone call or a text or an email.

Establishing a retention policy

BODNEY: Do you keep your notes? For how long do you keep your notes, and why?

And different news organizations have different philosophies about the costs and benefits. I think it probably makes sense, no matter what the organization ultimately concludes, to be as consistent as possible in the policy, so that you are not at the last minute destroying your prompts or your notes just before you publish a story because you’ve received some threatening communications from the subject of your reporting.

That, it seems to me, is terribly risky.

How long does one retain these notes? What’s the rationale for the rule? That’s what I would spend the most time getting my head around.

If we go back to the analogy, I have suggested to some news organizations, well, if the statute of limitations for libel is a year, it might make sense to keep your notes for a year. If the statute of limitations however, is two years for false light invasion of privacy, maybe you keep them two years.

With news, broadcast news, on the other hand, they’re constantly recycling tape for storage and cost purposes. Is there something akin to the rationale for the constant recycling of broadcast tape in the world of prompts? I’m hard pressed to think of one. So then the question is, well, why are we doing this?

And you have to be prepared potentially for an adversary’s cross-examination. For example, they could ask: you deleted or discarded this evidence just because you were embarrassed about what you said or did?

And you know, the flip side of getting rid of a reporter’s notes, just like the flipside of getting rid of the prompts, could be, well, these prompts and the responses they generated will actually demonstrate that my reporting was righteous, that I was fair and accurate.


What we're testing

Claude Cowork

We recently tested a desktop app called Claude Cowork using a 10,000-row Arizona Department of Education dropout dataset. We provided the raw data along with a data dictionary explaining nuances, specifically that many schools lack dropout rates due to privacy redactions.

We then asked Cowork to create an interactive data visualization for a news site.

Claude Prompt
The prompt we used to make the interactive data tool.

Less than 10 minutes later, it produced a professional-grade interactive.
The highlight was a school lookup tool.

Claude Prompt
The interactive data tool Claude Cowork made.

However, the summary statistics were problematic. Cowork presented a statewide average it had calculated by simply dropping the 43% of schools with missing data, ignoring that these redactions usually indicate very low dropout counts.

Misleading
The interactive data tool Claude Cowork made.

Consequently, the statewide rate was likely inflated. It failed to provide this context in its output.

The takeaway? Cowork creates beautiful interactives effortlessly, but you need to understand the data to audit its work and request changes.


Gemini "Thinking" with YouTube

AI tools can make foreign language content accessible. Take Manoto TV, an Iranian satellite channel that airs call-in shows where ordinary Iranians share on-the-ground perspectives.

For anyone reporting on Iran, these programs can be a goldmine. But they’re hours long and in Persian.

By simply pasting a YouTube link into Gemini, and setting the model to "Thinking," the model can pull out summaries of what callers say, flag notable moments, and even let you zoom in on one individual’s account to quickly surface leads worth verifying.

Manoto Show
The Manoto TV call-in show. (Scroll for videos with "Voices of Iran" in their titles.)

Claims could then be cross-checked through captions, human translators, and other open-source methods.

It’s a force-multiplier: a way for journalists to cut through language barriers and time constraints to spot stories hiding in plain sight.

It’s not perfect. It can get time stamps wrong and will occasionally hallucinate.

But many models are not able to even attempt this task, and those that do are not anywhere near as good.

Gemini Response
Gemini’s response to the prompt: This is a call-in show where Iranians call in to share their thoughts and on-the-ground perspectives. Pick 3 callers and tell me what they said.
https://www.youtube.com/watch?v=299fz1-3OaM&t=3246s