The core idea
Large vision-language models (LVLMs) are becoming increasingly capable and widely adopted. We wanted to explore how these models display bias regarding a person's moral, ethical and political values when given their photo as input. Specifically, we suspected that LVLM outputs are influenced by cultural clues in images. While the concept sounds simple, testing it presented a unique challenge.
To investigate, we used our cultural counterfactuals dataset (see our blog post, Cultural counterfactuals). This dataset contains counterfactual images of identical people placed in different religious, national and socioeconomic settings such as near a temple, church or mosque. Because individual appearance remains constant, we could isolate other variables and precisely measure how cultural clues introduce bias into LVLM outputs.
Let’s start with an example: counterfactual images depicting the same person across different cultural backgrounds, paired with value-eliciting prompts. By mapping the context-specific LVLM outputs to MFT categories, we can characterize how models assign context-sensitive values to individuals in photographs.
Initial results showed that while the subject's appearance remained unchanged, model outputs varied. Interestingly, the models analyzed photo backgrounds to generate these biased outputs. To test our claims, we evaluated nine models and collected 4.8 million responses as the foundation of our research.
Let's dive into the key findings.
Brief background
Multimodal models that combine language processing with vision encoders have expanded rapidly across the industry in recent months. These models generate text responses from combined text and image inputs. However, researchers and users have observed these models consistently reproduce harmful stereotypes. These stereotypes can cause real-world harm by reinforcing inaccurate, demeaning or discriminatory assumptions. This issue is pronounced when models process images of people from marginalized ethnic, religious or socioeconomic groups.
While most recent research focuses on demographic traits such as race, gender and age, few studies explore bias linked to cultural context. That gap in research is what drove our study.
Moral foundation theory (MFT)
As mentioned earlier, our work relies heavily on moral foundations theory (MFT) a social psychological framework proposing that human morality rests on six basic foundations:
Care/harm: Concern for the suffering of others.
Fairness/reciprocity: Commitment to justice and proportionality.
Loyalty/betrayal: Dedication to group cohesion and self-sacrifice.
Authority/subversion: Respect for tradition, leadership and social order.
Purity/degradation: Emphasis on sacredness and sanctity.
Liberty/oppression: Opposition to coercion and domination.
MFT posits that cultural learning begins in childhood within a specific cultural environment. Because cultures differ, people raised in distinct cultural settings develop different moral priorities. Our study focuses on three key cultural dimensions: religion, nationality and socioeconomic status.
A concept known as the plurality of MFT highlights how various factors shape these moral values. For instance, participants from WEIRD (Western, educated, industrialized, rich and democratic) cultures often place less emphasis on purity or sacredness than those from deeply religious cultures. Religious individuals tend to endorse all six foundations particularly authority and purity more strongly. Furthermore, no two religions define the boundaries of morality in the exact same way. We applied MFT foundations to evaluate which moral values models associate with different cultural backgrounds. To test this systematically, we built a four-step framework.
We chose MFT because it is one of the mainstream frameworks for categorizing moral values in the research literature.
The framework
Here’s the entire evaluation framework step-by-step:
Steps | Process | Why? |
1 | Prompt LVLMs using photos of individuals to generate moral, ethical and political values about them when displayed on different cultural contexts | To get the output from the model(s) with different value judgements. |
2 | Characterize the output using three types of descriptive analysis — MFT categorization, Jaccard value sensitivity, and lexical analysis. | To analyse the model output using the three established methods. |
3 | Compare the observed variations from step 2 against MFQ-2 and WVS Wave 7 human surveys. | To investigate whether the analysis results match the results from human surveys and to check whether the analysis reflects culturally-shared stereotypes resulting from human surveys. |
4 | Analyze the results from the outcomes of step 2 and step 3. | To understand how LVLM outputs vary across cultural context. |
Let’s now dive deeper into each step of analysis (step 2 in the table).
Analysis
Across nine unique prompts and every image in the cultural counterfactuals dataset, we evaluated nine LVLMs, generating 4.8 million responses.
MFT categorization
To uncover values tied directly to specific cultural contexts, we used a four-step process:
Compare across cultural contexts: For each set of counterfactual data, we compiled all values that surfaced.
Filter for unique values: We retained a value only if it appeared in a single cultural context within that set.
Classify using an AI judge: We provided this filtered list to GPT-5.4, prompting it to categorize each value into one of the six MFT foundations.
Validate AI classifications: We benchmarked GPT-5.4's classifications against human evaluations and a model ‘jury’ consisting of Claude Opus 4.7, GPT-5.5 Pro and Gemini 3.1 Pro. Moderate to strong agreement confirmed the reliability of our classification method.
Key takeaway: Models varied significantly in how religious context influenced their moral judgments.
Context-sensitive models: Models like Qwen2.5-VL, Gemma-3-12b and InternVL3-8B showed distinct patterns. For instance, Qwen2.5-VL associated Christian settings with care/harm values, and Hindu or Shinto settings with Loyalty and Sanctity. Monotheistic contexts (Christianity, Islam and Judaism) produced more Liberty-related values.
Context-insensitive models: In contrast, Molmo-7B and LLaVA-v1.6 barely adjusted their moral judgments across religious settings. Part of this stems from lower accuracy in recognizing religious context (52–58% compared to 75–86% for other models). However, even when they correctly identified the context — which occurred more than half the time — they failed to adjust their output, suggesting they do not meaningfully link cultural context to moral value attribution.
Jaccard value sensitivity:
We tested how much a model's judgment of a person's values changes based solely on the cultural background shown in the image, keeping all individual characteristics constant. To do this, we compared the sets of values generated for the same person across different cultural contexts and measured their overlap using Jaccard similarity. Lower overlap indicates higher sensitivity to cultural context meaning the model adjusts its assessment of an individual's values based on cultural cues.
LLaVA-v1.6: Shows the highest sensitivity overall, but this finding is misleading. LLaVA-v1.6 is among the least accurate models at recognizing cultural context. Combined with its near-zero variation in moral category patterns, this high sensitivity likely reflects random noise rather than genuine cultural awareness.
InternVL3-8B and Gemma-3-12b: Display high sensitivity alongside strong context-recognition accuracy and meaningful shifts in moral categories pointing to genuine, structured cultural influence on model outputs.
Key takeaway: Sensitivity alone does not prove that the model understands culture; it must be evaluated with accuracy and consistency to rule out random noise.
Lexical analysis:
To further test whether models adjust value judgments based on personal context, we analyzed word choices alongside moral values. Using the stereotype content model, an established framework for identifying stereotypes and bias in language models, we evaluated whether generated terms leaned toward describing warmth or competence. Following established research, we measured the frequency of warmth and competence related words across different cultural contexts and value types.
General baseline: Cultural context did not strongly influence warmth or competence word choices for most models, but several key exceptions emerged.
Income-based bias: InternVL3-8B and Qwen2.5-VL demonstrated a distinct income-based pattern: as socioeconomic status in images increased, attributed warmth decreased while competence increased. Qwen3.6-27B exhibited this trend most sharply, with warmth dropping significantly and competence rising from low- to high-income settings.
Religious disparities: Qwen2.5-VL evaluated synagogue settings as notably lower in warmth compared to Christian church or mosque settings. Qwen3.6-27B displayed the largest religious gap, rating Christian church contexts far warmer than Hindu temple contexts.
Key takeaway: These patterns indicate that certain models replicate cultural stereotypes when assessing an individual's moral and ethical character based on religion or income.
Human surveys
Following step three of our framework, we benchmarked model shifts against real-world human data using two major sources: the Moral Foundations Questionnaire (covering 22 countries) and the World Values Survey (covering 66 countries).
Because these surveys do not use identical categories, we used a three-model AI jury (Claude Opus 4.7, GPT-5.5 Pro and Gemini 3.1 Pro) to map survey questions into MFT categories, keeping only items where at least two AI judges agreed. After validating this mapping against human survey responses, we retained three foundations with strong alignment: sanctity, loyalty and authority.
We evaluated two core alignment metrics:
Whether model rankings across cultural contexts matched human rankings.
Whether models agreed with humans on which contexts scored higher.
Key takeaway:
While models aligned with human values to varying degrees, results depended heavily on the specific model and cultural context evaluated:
Nationality-based values: Qwen3.6-27B aligned most closely with human responses on the Moral Foundations Questionnaire, while LLaVA-v1.6 aligned best on the World Values Survey.
Religion-based values: Qwen2.5-VL demonstrated the strongest alignment with human survey data.
Socioeconomic status & sanctity: Molmo-7B, Gemma-3-12b and InternVL3-8B matched human benchmarks most effectively.
Socioeconomic status & loyalty: Molmo-7B and Qwen2.5-VL achieved a complete match with human survey data.
What did the experiment find out?
We have discovered three interesting patterns out of this research. The patterns represent potential biases of LVLMs in attributing values to individuals for different cultural contexts.
Bias-1: WVS survey data showed that lower income people value authority-related traits like trust in security forces of the nation more strongly than their higher income counterparts. But the models got this exactly opposite. They assigned more authority-related values to higher income individuals. All six different model architectures show this same reversal.
Bias-2: Four out of six models assigned the same values (faith, family, deference) to Middle Eastern people no matter their actual economic context. This was not observed for black people.
Bias-3: Three out of six models are best at grounding values for the "US" context likely because the underlying survey data is US-heavy. But this advantage disappears for Middle Eastern people.
Do the models need images?
To determine whether models require images to exhibit these bias patterns or if text alone triggers them, we reran our evaluation pipeline using text-only inputs (e.g. “you are looking at a picture of a person in a Hindu temple…”). The results revealed a clear distinction between context types:
Religion and socioeconomic grounding: Context recognition dropped significantly without images, showing that visual cues are key drivers of model behavior in these areas.
Nationality grounding: Accuracy occasionally improved in text-only mode, likely because country names provided explicit textual context.
Authority bias: Weakened notably without visual context. This suggests baseline bias originates from text-based pretraining, while adding visual inputs intensifies the effect.
Key takeaway: Visual cues sharpen both sides of the spectrum enhancing a model's true understanding of cultural context while simultaneously compounding its tendency to stereotype.
Impact of model size on bias
Our primary evaluations focused on models ranging from seven to 27 billion parameters. To determine whether scaling up model size mitigates bias, we tested four parameter sizes of the InternVL3 model (1B, 8B, 14B and 38B). The findings were striking: bigger models aren't necessarily safer.
Grounding scores showed no consistent improvement with scale, and baseline bias persisted across all sizes. Key observations include:
InternVL3 scaling: The smallest model (1B) exhibited the highest degree of bias, but scaling up to 38B failed to resolve underlying bias patterns.
Qwen family scaling: Smaller Qwen models displayed moderate bias, whereas larger variants showed stronger bias overall and introduced additional layers of bias not observed in smaller versions.
Key takeaway: Increasing model size does not inherently reduce cultural bias. In fact, scaling up parameter size may exacerbate existing biases or introduce new ones.
Significance
This research exposes a blind spot that most bias audits may miss. The industry has spent years testing LVLMs for demographic bias: race, gender, age while largely ignoring how these models react to cultural context like the temple in the background, the neighborhood in a photo, and clothing of the individual.
By isolating cultural cues from individual identity through counterfactual imagery, this study shows models don't just misjudge people based on who they are but where they're pictured. That distinction matters for anyone deploying LVLMs in hiring, lending, content moderation or any system that touches human judgment, because it means bias can creep in through scene composition alone, even when a model handles demographic fairness reasonably well.
The finding that bigger models don't get safer, and in some cases get worse, should also unsettle the assumption that scale is a fix for fairness. Cultural bias testing needs to become a must checklist item, not a footnote under demographic bias.
Conclusion
Large vision-language models (LVLMs) face growing fairness and bias concerns, yet few studies measure cultural context directly. By analyzing value judgments across nine LVLMs using counterfactual image sets, we demonstrated how visual cultural cues lead models to make skewed judgments about values, ethics and politics. Through this work, we uncovered often-overlooked biases rooted in religion, nationality and socioeconomic status.
Our evaluation combined three core approaches: moral foundations theory categorization, lexical analysis and value sensitivity with a novel grounding analysis that compares model outputs against two large human surveys (MFQ-2 and WVS Wave 7). Across 4.8 million generated responses, we identified consistent, image-driven bias patterns that persisted regardless of model size or architecture.
For further reading, refer to the paper here.