The Health Insight Survey That Couldn’t Handle Too Much Insight

Dark editorial illustration of a guillotine cutting through a Health Insight Survey response at the 250-character limit. A Grim Reaper operates the blade while the severed lower section, containing discarded feedback, is stamped “EXCESS INSIGHT”.

At the beginning of July I signed up to the National Health Insight Survey.

I’m not ordinarily the sort of person who fills in surveys for “fun”, but when it comes to the NHS I think it’s important to air an opinion. The Health Insight Survey is conducted by the Office for National Statistics and funded by NHS England, and its entire purpose is to give people the opportunity to regularly feed back on their experience of NHS services.

On the face of it, a perfectly sensible idea.

Ask people what happened. Find out where things work and where they don’t. Use that evidence to make better decisions about services used by millions of people.

However…

I lasted approximately five minutes before losing my shit in my favourite WhatsApp group, made up of some of my most trusted peers of the industry, where we trash-talk everything from research methods to cocktails.

My opening contribution to the discussion was, as usual, remarkably restrained:

“you want feedback? i’ll give you feedback you f*ckers! this is for an NHS survey btw…”

And in that rage I attached a screenshot of the feedback I’d just submitted to the survey itself, when they asked if I’d like to give feedback about the survey:

“If you want genuine rich insights from these surveys stop putting stupidly restrictive character limits on free text response boxes. You will get nothing of value unless you allow people to articulate the issues fully. You are biasing your own results.”

Screenshot of the Health Insight Survey feedback form containing criticism of restrictive free-text character limits, shared in a WhatsApp message saying, “you want feedback? i’ll give you feedback you f*ckers!”
My entirely measured response after discovering the Health Insight Survey wanted qualitative feedback, just not very much of it.

One group member immediately replied:

WhatsApp reply reading, “So you first typed whatever feedback, but it didn't fit because of the character lim,” deliberately cutting off the word “limit”.
Peer review was swift. And yes, “character lim” was entirely deliberate.

“So you first typed whatever feedback, but it didn’t fit because of the character lim”

Yes. He cut himself off deliberately. I would expect nothing less.

Another was less concerned about research validity and more disappointed by my complete failure to exploit the commercial opportunity:

“You forgot the call to action, like your calendly link”

Apparently even incandescent rage needs a funnel.

WhatsApp reply reading, “You forgot the call to action, like your calendly link,” in response to a message ranting about the Health Insight Survey.
Friend identified the real failure in my survey feedback: absolutely no attempt at lead generation.

Anyway, I moved on with my life. Or at least I thought I had.

Then the box came back.

Roughly a month later another Health Insight Survey arrived, and somewhere between the usual ambiguous questions and me trying to remember what I’d answered last time, my old nemesis appeared again.

The stupidly oddly large but deceptively restrictive free-text box, but this time in the ACTUAL survey responses they wanted me to give.

It has therefore taken me this long to recover sufficiently from the original rage to ask a slightly more useful question than the one going through my head at the time, which was basically: WTAF.

Because underneath the swearing there is actually a fairly serious research problem. If you know me, this is usually the case.

This time the question was: “What is the main reason why your experience was neither good nor poor?”

Great, I thought. Time to give some valuable feedback so they can improve their survey response insights.

So I answered as truthfully as I could:

Because they haven’t done anything good, and nor have they done anything bad. If you know how to work the system its ok, but speaking to a human in the first instance is still problematic. Not everyone is comfortable doing stuff online. There should

And there that sentence sat, suspended in time and never to be completed.

Why? Because as I kept typing, a big red warning appeared telling me I had hit the character limit.

The limit was 250 characters. Including spaces.

And this wasn’t some weird anomaly field where somebody had accidentally knocked a zero off. The ONS methodology for the survey explicitly says that all free-text questions in the Health Insight Survey have a 250-character limit. Enraged by the stupidity of limiting what I could say, I had another stab purely to vent:

Because they haven’t done anything good, and nor have they done anything bad. If you know how to work the system its ok, but speaking to a human in the first instance is still problematic. Not everyone is comfortable doing stuff online. This box is too small, you can’t get decent qualitative open feedback when you limit the character limit. Duh.

The survey responded:

You have exceeded the character limit by 97 characters.

Full disclosure: the bit where I start complaining about the box itself was deliberate. Once I realised what was happening, I carried on typing because I wanted to see exactly how ridiculous the constraint was.

Health Insight Survey question asking for the main reason an experience was neither good nor poor, with a free-text response and a red warning stating, “You have exceeded the character limit by 97 characters.”
Asked to explain the experience. Allowed 250 characters, including spaces. My answer exceeded the limit by 97.

But strip away my deliberate messing about at the end and the substantive answer I was trying to give, my character count was already at 236.

That left me 14.

Fourteen characters to qualify something, add an example, explain a second problem or say why any of it might matter.

And there was already quite a lot going on in those 236 characters.

I wasn’t saying my experience had been vaguely average. Nothing particularly good had happened and nothing particularly bad had happened either. That is different, although it also turns out to be surprisingly difficult to explain why something is neither good nor poor.

How do you justify the immediate response of:

“It’s meh, but not meh enough for me to say why either way”?

After thinking about it, part of the reason was that if you already know how to navigate the system you can generally get where you need to go.

But getting to a human in the first instance can still be problematic.

And that matters because digital access does not work equally well for everybody. This is the NHS. Its audience is essentially everyone, which means differences in digital confidence, access, ability and preference aren’t edge cases you can politely wave away because an online route exists.

Interestingly, ONS’s own analysis of Health Insight Survey responses finds both sides of exactly this problem: some people describe digitalisation positively, while others report difficulty contacting their GP practice online or by telephone. Human experience remaining stubbornly contradictory. How inconvenient.

The points in my answer were related, but they weren’t the same observation.

  • One was about the overall experience.
  • One was about existing knowledge of how the system works.
  • One was about access to a person.
  • One was about the assumption that providing a digital route means everyone can use it equally well.

Potentially different problems. Potentially different people affected. Potentially different things worth investigating.

But I had been asked for my main reason, and then given 250 characters to explain it.

Which means something fairly important has happened before a researcher has even seen my response.

I have been made to edit the evidence for them.

I have to decide what survives.

  • Do I remove the bit about knowing how to work the system?
  • Do I remove the human contact problem?
  • Do I remove the point about people who aren’t comfortable online?
  • Which part of my experience would you like me to pretend didn’t happen so the rest fits neatly into your database field?

That isn’t a neutral act of collection.

The survey has already influenced the evidence by defining how much of my explanation it is prepared to receive.

And this is where it gets particularly idiotic.

ONS explains why it introduced open-ended questions into the Health Insight Survey. Closed questions tell them how people rated their experience, while the open questions let people explain the reasons behind those ratings.

In their words: “This allows us to gain more insight into the NHS patient experience.

Excellent. I have insight. Quite a lot of it, apparently. Unfortunately some of it starts at character 251.

We tend to talk about surveys as though there is a nice clean sequence involved. Something happens to a person, you ask them about it, they tell you what happened and then you analyse what they said. There are decisions all the way through that chain.

  • The wording of the question changes what people think you’re asking.
  • The answer options change what they can tell you.
  • The order of the questions can affect responses.
  • And a character limit determines how much of somebody’s explanation is permitted to exist in the dataset at all.

The interface isn’t just carrying the research.

The interface is part of the research instrument.

If it changes what people are able to tell you, it changes the evidence you collect.

The missing bits are gone for good. Poof. Rabbit in a hat.

An analyst looking at my eventual response doesn’t see the sentence I deleted. They don’t see the caveat I couldn’t fit. They don’t see the second problem I decided was less important than the first.

They see whatever survived the 250-character purge and quite reasonably treat that as the answer I wanted to give.

It isn’t. It’s what managed to fit.

Humans are inconvenient.

Now, there used to be at least a practical argument for keeping qualitative feedback relatively constrained.

Give thousands of us an open-text box and we’ll fill it with thousands of different descriptions, examples, caveats, complaints, weird edge cases and rambling stories. Then somebody has to sit down and make sense of the bloody lot.

Reading it takes time. Coding it takes time. Finding themes takes time. Working out whether twenty people are describing the same underlying problem using completely different language takes time.

Qualitative data is messy because people are messy.

So you can understand why organisations have historically been tempted to make the mess smaller. Tick boxes. Scales. Categories. Tiny text fields. Nice little bits of humanity that fit conveniently into whatever system is waiting at the other end.

There is a slight problem with using that as a defence here though.

We don’t actually know why the Health Insight Survey has a 250-character limit. It might be technical. It might be operational. It might be a design decision. It might have a perfectly sensible explanation sitting somewhere I haven’t found. So I’m not going to invent one for them.

What we do know from the published methodology is what happens afterwards.

ONS doesn’t manually code every free-text response. For its published analysis of GP experiences, researchers randomly selected subsets of the responses and used codebook thematic analysis until they reached data saturation — the point where analysing more responses stopped producing new patterns. They coded 1,276 responses concerning positive experiences and 928 concerning negative experiences from much larger response pools.

This is interesting, because the assumption that collecting richer qualitative responses necessarily means somebody has to manually read every single word from every single person doesn’t apply here anyway.

They already have a research method for dealing with volume.

And now, of course, there is AI.

Before everyone runs screaming towards the nearest LLM with a CSV file, no, I am not suggesting we collect 50,000 enormous responses, shovel them into an AI, ask “what do people think?” and put whatever comes back into a PowerPoint labelled “RESEARCH”.

That would merely introduce a shiny new way of mangling the evidence, and sadly we’re already rather good at that.

But it is 2026, and the argument that we need to force human experiences into tiny pieces simply because unstructured language is difficult to work with is becoming increasingly boring.

AI can already help researchers navigate large amounts of qualitative material. It can assist with finding recurring themes, keywords and candidate patterns, searching responses for particular experiences and getting an initial view of a large body of text much faster than was previously possible.

That doesn’t make it a researcher.

A 2025 study in Scientific Reports compared GPT-4o’s qualitative analysis with human analysis and found useful applications for identifying themes, keywords and basic narrative, but also problems with hallucinated or altered quotations, context and the depth of analysis. The researchers’ conclusion was essentially: useful assistant, absolutely not a replacement for experienced qualitative researchers.

Good. That’s roughly where I am too. Use it to help with the work.

But for the love of whatever deity you identify with, or simply because you don’t want to be an idiot, don’t outsource judgement to it.

Keep the raw responses. Read actual verbatims. Check whether the themes it gives you exist. Look at the outliers. Go back to the evidence. Triangulate with other sources.

Basically all the boring research discipline we were supposed to be doing anyway. The important bit happens before any of that.

Collect the evidence first.

Because that part isn’t reversible.

If somebody gives me 1,000 characters describing an experience, I can reduce that later. I can categorise it. Code it. Search it. Summarise it. Ask questions of it. Give a machine a first pass and then go back to the original to check whether the machine has confidently made something up.

If I only allow them 250 characters in the first place, I cannot recover the rest from the ether later.

And that, thinking about it, is the part of all of this that annoys me most.

We’re living through a moment where the cost and difficulty of working with messy human language is dropping dramatically, yet we’re still designing research mechanisms that require humans to make their experiences conveniently machine-shaped before we’ve even collected them.

We throw away richness at source.
Then we analyse what’s left.
Then we call the result insight.

The chain looks something like this: Reality → Interface → Evidence → Analysis → Decision

Add AI and there is another layer: Reality → Interface → Evidence → AI synthesis → Human interpretation → Decision

There are already enough opportunities in that chain for nuance to disappear without deliberately chopping bits off at the entrance before the show has even started.

Healthcare doesn’t arrive in tidy little boxes.

And let’s also not forget what we’re talking about here.

Healthcare experiences are almost designed to resist neat little answers.

  • The medical care can be excellent while getting an appointment is awful.
  • The digital service can be brilliant for one person and an obstacle for another.
  • Someone can be enormously grateful for the treatment they received while being furious about what they had to go through to receive it.
  • Something can work perfectly once you understand how the system operates while being bewildering to somebody encountering it for the first time.

Even ONS’s own analysis points out that qualitative importance comes from depth, context and meaning, rather than simply how frequently something appears.

Which is, frankly, doing a fair amount of work for my argument.

Experiences like these don’t arrive pre-categorised.

They are messy. Conditional. Contradictory. Sometimes you need more than one reason to explain them.

Which, presumably, is why you ask people to explain.

So let them.

The box was never really the problem.

My objection was never really that I couldn’t fit my little rant into the box. That bit was just me rage-slapping my keyboard.

It’s that the survey asks for insight and then places a hard boundary on how much of someone’s explanation it is prepared to collect.

We spend a lot of time worrying about bias in research participants. We should probably spend at least as much time wondering what the research instrument itself has done to the evidence before anybody starts analysing it.

You can always summarise evidence later.

You cannot analyse what you refused to collect.

And if the purpose of open-ended questions is, in ONS’s own words, to “gain more insight into the NHS patient experience”, perhaps the first step is to stop rationing the insight.

Apparently it’s very valuable.
Just not after character 250.

Abi Hough

Abi Hough is the founder of Corpus and a UX, experimentation and accessibility strategist. Her work examines the gap between what organisations claim, what their systems communicate and what people ultimately understand.

Talk about how this applies in your organisation.

If a field note resonates and you want to talk about how the same patterns are showing up where you work, a conversation can help.
Typical first conversations last 45 to 60 minutes and focus on understanding your current situation and constraints.
Upstream optimisation for zero click and AI search.
Contact
[email protected]

For enquiries about the Corpus Diagnostic, speaking or other work.
We Are Corpus is an Abi Hough-led consultancy and the trading name of UU3 Ltd, registered in England and Wales. Company number 06272638.