The multilingual problem is harder than it looks
Most feedback analysis tools are built and benchmarked on English text. If a meaningful fraction of your customers write in other languages, you are typically using a tool that was not designed for your actual data — and accuracy will be lower than you might assume.
The multilingual problem has several dimensions. Sentiment lexicons are language-specific. Keyword lists and category examples need to be in the languages of your feedback to match. Tokenization and morphology work differently across languages. And the distribution of good multilingual models varies — some languages have much better coverage than others in existing model training data.
There is no configuration that makes multilingual analysis as seamless as English-only analysis. But there are practical approaches that work well enough for the most common use cases.
What VADER can and cannot do across languages
VADER is an English lexicon. Its sentiment scores are derived from English words and English grammatical rules. Feeding non-English text to VADER produces unreliable results:
- Unknown tokens — words not in the English lexicon score as neutral. A Spanish negative sentence full of Spanish words scores near zero (neutral) because VADER has no opinion about any of them.
- False positives from cognates — Spanish, French, and Italian share many words with English cognates. "Fatal error" in French reads as English "fatal error" to VADER. This sometimes produces accidental correct scores, but inconsistently.
- No grammatical rule application — VADER's negation and modifier rules are built for English grammar. They do not apply to languages with different syntactic structures.
If a significant fraction of your feedback is non-English, VADER's sentiment scores for those items will be close to neutral regardless of actual content. You can filter your VADER-based sentiment dashboards to English items, or treat the neutral score on non-English items as a signal that sentiment was not computed, rather than that the feedback was genuinely neutral.
TF-IDF clustering across languages
TF-IDF clustering is language-agnostic in the sense that it works on tokens regardless of language. The practical problem is that it will create language-segregated clusters: Spanish feedback will cluster with other Spanish feedback about the same topic, but that Spanish cluster will be separate from the English cluster about the same topic — because the vocabulary is different.
For a multilingual corpus, this means your topic clusters reflect language as much as they reflect theme. A single problem reported by English-speaking and Spanish-speaking customers will appear as two separate clusters rather than one. This is not wrong — it is an accurate reflection of the vocabulary distance — but it means you need to be aware that similar-sized clusters in different languages may represent the same underlying issue.
If you want cross-language topic coherence, the practical options are either to translate all feedback to a common language before clustering, or to use a multilingual embedding model that maps text in different languages to a shared vector space before clustering. Rereflect's current TF-IDF implementation does not handle this automatically.
Using an LLM for multilingual categorization
The most practical path to multilingual feedback categorization is a language model with strong multilingual capability. Most frontier hosted models (OpenAI, Anthropic, Google) handle a broad range of languages well, including common European languages and major Asian languages. Less common languages may have significantly lower accuracy.
When using an LLM for multilingual classification, a few configuration choices matter:
- Write category descriptions in English — most multilingual models have their strongest reasoning in English. Category descriptions in English work well even when the feedback text is in another language.
- Instruct the model to classify regardless of language — explicitly tell the model in the prompt that feedback may be in any language and that it should classify based on meaning, not language of origin.
- Check coverage for your specific languages — if you have significant feedback volume in a language you are not sure the model handles well, test it against a manually labeled sample before relying on it.
- Local models vary widely — Ollama and other local model runners offer models with different language coverage. A model like Mistral or Llama 3 may handle major European languages reasonably well; coverage for less common languages is usually weaker.
A practical configuration for mixed-language feedback
If your feedback is primarily English with some non-English items, the simplest approach is to let VADER handle English items and flag non-English items for LLM-based sentiment as well as categorization. Most LLMs can detect the language of an item and adjust accordingly.
If your feedback is substantially multilingual — say, 30% or more non-English — it is worth investing in a multilingual embedding approach for clustering, or committing to LLM-based processing for all categorization and sentiment tasks. The keyword pre-filter will be less effective because your keyword lists are likely English-only; relying more heavily on the LLM layer compensates.
The honest position is that multilingual feedback analysis is more difficult than English-only, and the gap in accuracy is real. The most important thing is to know which items were analyzed in which way, so you can interpret your dashboards accordingly and not treat a high neutral-sentiment rate in non-English items as a product signal when it is actually an analysis gap.