Better Method.ai
Recipe 02

Turning 400 evaluation comments into something you can act on

The open text box is where the real feedback lives and where it goes to die. This gets you themes, counts, and the quotes that matter, without quietly softening the criticism.

What you’re making

A summary of your open text responses: named themes, how many comments sit under each one, real quotes you can use, and the handful of comments that a theme count would otherwise bury.

This is probably the most useful thing on this whole site. It is also the one with the most ways to go wrong. Both of those are worth taking seriously.

What to have open

Your export, with the ratings columns removed and the names taken out. One comment per line.

If you have last cycle’s theme list, keep it handy. Comparing this cycle against the last one is more useful than starting fresh, and it’s a question the tool answers well.

The prompt

Prompt: copy and adapt
I have [number] open text responses from an evaluation of [activity name and format]. Respondents were [describe the learners].

I want an analysis I can defend, not a summary that makes us look good.

Rules:
- Work only from the comments I paste. Do not infer, extrapolate, or add anything that isn't in the text.
- Do not soften criticism. If comments are blunt, the theme name should be blunt.
- Keep the small stuff. A point raised by 4 people out of 200 is not noise if it names a specific problem. Report those separately instead of dropping them.

Give me:

1. THEMES. For each one: a short plain name, how many comments fall under it, one sentence on what people are actually saying, and 2 real quotes (exact, uncorrected, with any bracketed placeholders left as they are).

2. THE UNCOMFORTABLE SECTION. The three most critical points made, stated as directly as the respondents stated them. Do not blend them into something gentler.

3. RAISED BY ONLY A FEW PEOPLE. Comments naming a concrete, fixable problem that hardly anyone mentioned. These are the ones a theme count buries.

4. ACTIONABLE OR NOT. Split the feedback into what we could actually change next time, and what is context we can't change (parking, weather, the subject matter itself).

5. WHAT I DIDN'T ASK. Anything notable in these comments that none of the above captured.

Do not give me an overall satisfaction judgment or a score. I have the numbers separately.

COMMENTS:
[paste your anonymized comments, one per line]

What good output looks like

The sign of a good pass is that it makes you slightly uncomfortable. If it reads like something you’d happily forward to leadership without editing, something has gone wrong. Models lean hard toward polite paraphrase, and the “do not soften” instruction is fighting a real habit.

Good output looks like this:

Theme: the case discussion got rushed to make room for the last two talks. 31 comments. People aren’t saying the cases were bad. They’re saying the format promised interaction and then didn’t deliver it. “We got 8 minutes for the case that was supposed to be the whole point.” “Third year in a row the discussion gets cut. Just build the agenda honestly.”

Now compare that with what you get without the constraint: “Several participants offered suggestions regarding time allocation.” Both are technically true. Only one of them tells you what to do differently.

How to check it

Spot check the counts. Take the two biggest themes and search your source file for the words you’d expect to find. Counts drift, because models approximate when they tally long lists. Don’t put a number in a report you haven’t sanity checked. If the exact counts matter, ask for the analysis in two halves and add them up yourself.

Check every quote. Search your source file for each quote in the output. Models paraphrase while claiming to quote. A quote in an outcomes report that nobody actually wrote is a credibility problem, and it’s easy to catch: search, don’t skim.

Read the raw comments yourself anyway. Skim all of them, quickly. You’re not checking the analysis. You’re checking whether anything in there needs a human response today: a safeguarding concern, a serious allegation about a faculty member, someone who sounds genuinely distressed. No summarizing tool will flag that for you, because it has no idea that escalation is a thing.

Where it goes wrong

It softens the criticism. Left alone it rounds everything toward neutral. This is the failure that matters most, because it produces a report that is pleasant, defensible and useless.

It merges themes that don’t belong together. Two different complaints get combined because they share a word. “The room was too cold” and “the pace felt cold and impersonal” have landed in the same theme before now. Read the quotes under each theme and check they belong.

The counts drift. Approximate tallies stated with total confidence. Assume every number needs checking.

It can invent structure when you give it too little. A small comment set may not support stable themes. Review the comments yourself and treat any generated grouping as a draft, not a finding.

Copilot
Microsoft 365 Copilot can work with supported Excel files and permitted SharePoint or OneDrive content. Remove identifiers in a separate working copy first, and confirm that your license and policy allow the source file to be used.
ChatGPT
In an approved workspace, upload the screened working file rather than pasting hundreds of lines. Verify every count and quote independently before it enters a report.
Claude
A paid Claude Project can keep a screened file and standing analysis instructions together. Verify every theme, count and quote against the source.

The part nobody tells you

Once you’ve done this twice, the interesting question stops being what did people say and becomes what has changed since last time.

Paste last cycle’s theme list alongside this cycle’s comments and ask which themes have been resolved, which have stuck around, and which are new. That comparison is what turns evaluation data into a story about your program, and it’s the thing your accreditation report actually wants.