Why the prompt matters
A language model's output is a function of its weights and its input. You cannot change the weights — that is the model's training — but you entirely control the input. For a feedback categorization task, the input is the prompt, and the quality of the categorization depends on how much useful context that prompt contains.
This is not a theoretical point. The same feedback item sent to the same model with two different prompts can produce meaningfully different category assignments. A prompt that describes categories vaguely produces vague assignments. A prompt that gives the model clear, distinct category definitions produces cleaner assignments. The effort you put into your taxonomy descriptions shows up directly in categorization quality.
The anatomy of a feedback categorization prompt
A well-structured categorization prompt for feedback analysis contains several components:
- Task framing — a clear statement of what the model is being asked to do: classify this feedback item into one or more of the following categories.
- Category definitions — the full set of categories with their names and descriptions. This is where your taxonomy configuration is used; whatever you have written as the description for each category gets included here.
- The feedback text — the actual item being classified, clearly delimited from the surrounding prompt.
- Output format specification — explicit instructions for how to return the result (JSON with category IDs, comma-separated labels, structured fields), so the output can be parsed reliably.
- Edge case handling — instructions for what to do when no category fits, when multiple categories apply, or when the item is too ambiguous to classify confidently.
Getting the output format specification right matters as much as the category definitions. An LLM that produces well-reasoned classification but returns it in a format the parser does not expect produces useless output. Structured output modes (available in most hosted APIs) help, but even without them, a clear and specific format instruction dramatically reduces parsing failures.
Writing category descriptions that work
Your category descriptions are the most important variable in classification quality. A few principles hold across different feedback domains and model choices:
- Describe the category, do not just name it — "Performance" tells the model almost nothing. "Slowness, loading delays, timeouts, and high latency in any part of the product" gives it something concrete to match against.
- Include examples of the vocabulary your customers actually use — if your customers say "laggy" and "stuck," mention those. The model needs to connect your category definition to the language of your feedback.
- Make categories mutually exclusive where possible — if two categories overlap significantly, the model will inconsistently assign items that could fit either. Draw the boundary explicitly: "This category covers X but NOT Y — items about Y belong in the Z category."
- Describe what is excluded, not just what is included — a category definition that only covers what belongs generates more false positives than one that also clarifies what does not belong.
- Avoid jargon the model may not know — internal product names or proprietary terminology may not be in the model's training data. If a category is defined around a product feature with a proprietary name, describe what the feature does, not just what it is called.
Token budget and cost
Every character in the prompt is a token that costs money (on a hosted API) and adds latency. Prompt design for feedback categorization involves a tradeoff between thoroughness and cost.
The category definitions are repeated for every item classified. A taxonomy with ten detailed category descriptions might add several hundred tokens to every call. At scale — thousands of items per week — that adds up. The practical strategies:
- Keep descriptions precise, not exhaustive — a well-targeted 40-word description often outperforms a rambling 200-word one, and costs a fraction as much.
- Avoid redundancy across categories — if the same phrase appears in multiple category descriptions, it is doing no work. Descriptions derive their value from distinctiveness.
- Consider the keyword pre-filter — Rereflect's keyword layer handles items that are obviously in one category, reserving LLM calls for ambiguous cases. This reduces your effective per-item token spend without sacrificing accuracy on the hard cases.
Iterating on prompts in practice
Prompt design is empirical, not theoretical. The way to know whether your category descriptions are working is to sample classified items, look at the ones that were miscategorized, and ask: what information would have prevented this error?
If the model assigned "performance" to an item about "billing latency" — and you have separate categories for performance and billing — the solution is usually to add language to the billing category description that explicitly includes payment-related slowness, and to add language to the performance category that excludes billing-related delays.
Make one change at a time, re-run on a small sample, and check whether the target cases improved without causing regressions elsewhere. Prompt iteration is quick — it does not require retraining anything — but it benefits from the same discipline as any other A/B-style comparison: change one variable, measure the effect.