Case Study

Running market research as an agent swarm

Three countries, 19 sourced briefs, and a standard strict enough to fail my own work.

8 min

3

Country studies

US, Thailand, India, run in parallel

19

Sourced briefs

One narrow mandate per agent

305

Unique cited sources

Inline link plus access date

28

Structured data tables

Competitor matrices, price ladders

The problem

A supplement brand with more than 20 SKUs wanted to enter the United States. Its founder had fully exited, which meant no audience, no creator channel and no inherited trust. The launch would have to win on catalogue selection, listing quality and compliance alone.

That is a research problem before it is a marketing problem. Which products, in which market, at what price, against whom, and what can we legally say about them. Doing it properly across three candidate markets is weeks of work for a team. I did not have a team.

The shape of the system

Instead of one agent asked to research a market, I built swarms of agents, each given a single narrow mandate and each producing exactly one structured, fully sourced document. Narrow mandates are the point: an agent asked to do one thing produces something you can check, and an agent asked to do everything produces something that sounds right.

1

Wave 1, run in parallel: competitor intelligence and demand or format analysis. Then a checkpoint.

2

Wave 2, run in parallel: catalogue fit, compliance screening, and voice of customer. Then a checkpoint.

3

Handoff: wave 1 produces a keyword pull list. The client runs it in their own marketplace tooling and drops the exports back in, so the demand numbers come from their account rather than my inference.

4

Wave 3: a synthesiser consolidates, but only after the real demand data has landed.

Three of these ran at once, one per market, with separate agent sets so findings could not bleed between countries.

The rules every agent had to follow

The output is only as good as the constraints. Six rules bound every agent in every swarm:

Source everything. Every factual claim carries an inline source link and an access date.

No verbatim dumping. Paraphrase, never copy more than roughly 25 consecutive words.

Recency bias. Prefer 2024 to 2026 data and flag anything older explicitly.

Confidence ratings. Mark each major finding High, Medium or Low.

An honesty section. Every file ends with "What I could not verify."

Stay in scope. One market per swarm, existing products as they are, no drifting into redesign.

The honesty section is the one that changes behaviour. An agent that has to write down what it could not establish stops quietly filling the gap.

What it produced

OutputDetail
19 research briefsCompetitor intelligence, demand and formats, catalogue fit, voice of customer, entry economics, regulatory screening
Catalogue prioritisation20+ SKUs scored down to a three-SKU launch wave, validated against marketplace demand data rather than internal opinion
Margin gateA unit-economics pack that must clear before any listing goes live, so go or no-go is a number
Regulatory screenTwo jurisdictions: US FDA, DSHEA and FTC substantiation, and Indian ASCI, Drugs and Magic Remedies Act, AYUSH, FSSAI and Legal Metrology

The regulatory pass blocked several proposed product names and marketing claims before they reached production. Finding those at listing review, after packaging is printed, is the expensive version of the same discovery.

Then I audited it, and it failed

A research programme that grades itself is worth very little, so I ran the programme against its own stated rules. It did not fully pass.

What the audit found

The mandatory "what I could not verify" rule had been followed in 16 of 19 files, not all 19. The rule had never actually been audited.

Access-date coverage reached about 39 percent of citations, against a rule that said every claim carries one. "Every claim is sourced" was the rule, not the achieved state.

One synthesis document contradicted its own underlying raw data on a market-size figure. Both numbers could not be true.

The catalogue-fit files carried almost no source links, because they are an inference layer rather than a research layer. That distinction had not been made explicit anywhere.

I published all four rather than quietly fixing the ones that were convenient. The most useful of them produced a rule I now apply everywhere: never source a narrative claim from a synthesis document. Synthesis layers editorialise about earlier waves. If a claim matters, go back to the raw file.

The review gate earned its keep

Every deliverable passes an adversarial review before it reaches the client, scored against fixed criteria, returning a pass or a fail with reasons. Nothing ships on a fail.

On one content deliverable it failed the first version on three separate axes: a causal narrative that the underlying data did not support, client de-anonymisation through verbatim SKU names, and a flat assertion that something was illegal when the honest answer was conditional. All three would have been embarrassing in front of the client. One of them would have been a real problem.

A later lesson from the same gate: fixing a review finding can introduce a new one. A corrected draft removed a wrong figure and replaced it with a number that matched nothing. Re-verify the replacement, not just the removal.

What transfers to product work

Narrow mandates beat broad ones. Scope an agent, or a person, to something checkable.

Make the standard explicit and then measure against it. A rule nobody audits is a preference.

Put a gate between the work and the customer, and let it fail things.

Distinguish research layers from inference layers, and label which one a document is.

Publish your own misses. It is the only thing that makes the passes credible.

Questions#

Is this just prompting a model repeatedly?

No. The value is in the constraints and the topology: narrow mandates, parallel waves with checkpoints between them, a human-supplied data handoff before synthesis, six binding output rules, and an adversarial gate at the end. Take those away and you get plausible text at volume, which is worse than nothing for a decision this expensive.

How do you know the sources are real?

Partly by rule, every claim carries an inline link and an access date, and partly by audit. The audit is how I found that access-date coverage was 39 percent rather than 100. A rule without a measurement is an intention.

Who is the client?

An Ayurvedic and natural-wellness supplement brand entering the United States. The engagement is live and confidential, so the brand, its products, its founder and its city are not named anywhere in this write-up.

Did the launch work?

Nothing has listed yet. The margin gate, formula confirmation, certificates of analysis and trademark screening are all still open. Scope, rigour and decisions are what I can claim here. Sales are not.

Romil Sharma

Product Manager · Growth · AI Product · Market Research

Eight years across market research, growth leadership and shipped AI products. Currently leading international market entry for an Ayurvedic and natural-wellness supplement brand, and building knowledge and evaluation systems under Regnor.

Back to the homepage →
© 2026 Romil Sharma. All rights reserved.|Privacy