The Evidence Base for Colour Psychology Claims in Design
Colour psychology claims often outrun the evidence. This page maps what research actually establishes about colour and behaviour, and where cultural variation limits generalisation.
What the Research Actually Says
Before a single hex code is chosen, the question is not whether colour matters but what the research actually says. Findings come from psychophysics, cognitive science, and cultural anthropology, each with its own methodological limits. A developer or marketer who can speak that language avoids the twin failures of the field: credulous acceptance of every agency claim and cynical dismissal of all of it. The hard line is drawn where the measurement is real. Contrast ratios have a formula. Colour meaning varies by culture and by context, and those variations are documented. Arousal and valence respond to hue, saturation, and lightness in ways that replicate. What does not replicate is the grand claim that a particular hue causes a particular behaviour in every viewer. That claim is assertion, not evidence.
Culture Changes What a Colour Means
The most reliable finding in the field is that colour meaning is contingent. Red scores positively in China, where it signals luck and prosperity, and negatively in parts of West Africa, where it signals danger and death. These are stable cultural linkages with documented histories. Berlin and Kay's 1969 work on basic colour terms established that the median language has five or six terms, and the maximum observed is eleven or twelve. The number of terms shapes what can be distinguished, but it does not fix what a colour means. The 2012 colour-in-context theory formalises this: the same hue carries different meaning depending on what it is attached to. A red stop sign, a red dress, and a red warning light are not the same stimulus, and no research says they are. The ecological validity of a laboratory swatch is the first thing to question.
The failure case is the brand that picks a colour because a competitor's palette analysis said it 'means trust'. That analysis is a post hoc narrative, not a finding. The empirical evidence for colour and behaviour is stronger for what is avoided than for what is preferred. Treated wood in garden furniture is chosen for its colour, not its preservation; the colour is the behaviour. When a designer needs a hue to signal 'correct' or 'safe', the research supports green, since that is a learned link, not an innate one. If the market is global, the link is not safe at all. The only ethical move is to test with the actual audience, in the actual context, under the actual lighting. Anything else is a guess with a citation attached.
Where the Lab Falls Short
Context Is the Whole Finding
Colour meaning research limits begin with the lab. Most studies present isolated swatches on a grey background under D50 lighting, which is the ISO 3664 standard for graphic technology. That is a reasonable controlled condition, but it is not a shopping aisle, a phone screen in sunlight, or a living room lit by a warm LED. The context effect is not a footnote; it is the entire finding. Elliot and Maier's colour-in-context theory, published in 2012, was a response to exactly this problem: the same red that impairs performance on a detail-oriented task in one study improves it in another, and the difference is what the red is attached to. The hue-performance effect in cognitive tasks is real, but it is small, and it is moderated by the task's relevance to the colour. Red on a proofreading task is a different stimulus from red on a math worksheet.
What Replicates and What Does Not
The replication crisis has hit this field harder than most. A well-cited study on the colour of a pill changing its perceived effect was followed by a series of failures to replicate the magnitude. The placebo effect of colour in medication is real in the sense that red and yellow pills are more likely to be perceived as stimulants, and blue and green as sedatives, but the effect sizes are modest and the clinical relevance is uncertain. What a designer can take from the research is the direction of the link, not its intensity. The claim that 62 to 90 percent of snap judgments are based on colour alone is a marketing statistic with no peer-reviewed source. It survives because it is quotable, not because it is true. The evidence supports a simpler claim: colour is one input among many, and it matters most when other attributes are matched.
Myths That Will Not Die
No Hue Has a Single Fixed Meaning
Colour psychology myths start with the idea that any hue has a single fixed meaning. The 'colour of the year' industrial complex depends on this fiction, and it is a fiction because the same red that raises heart rate in one study is used to sell lipstick in another. The arousal effect of red versus blue is consistently found, but it is a physiological response to wavelength, not a semantic one. Red is more arousing than blue, and the effect replicates. What does not replicate is the claim that red makes people hungry. That finding, when traced to its source, is a 1970s study of a single restaurant chain's sales, and it has not survived contact with a controlled experiment. The same is true for the claim that blue suppresses appetite: no one has been able to demonstrate it outside the lab.
Perceptual Effects You Can Rely On
The colour-weight link has better support. Dark colours are perceived as heavier than light colours, and the effect size is large, above d = 1.0. The colour-size link, where light objects appear larger than dark ones of equal size, is also robust. These two findings are the most useful in the design toolkit because they are perceptual, not cultural. The Bezold effect, where the colour appearance shifts with the outline colour, and simultaneous contrast, where a neutral grey appears tinted with the complementary hue of its surround, are both documented since the Munsell Book of Color of 1929. These are not myths; they are the mechanical realities of human vision. The myth is that they need a psychologist to explain them. They need a designer who has seen them happen and knows the conditions under which they will.
What People Actually Prefer
Colour preference evidence is more stable than the meaning claims. The ecological valence theory, which derives preferences from the average emotional response to objects of that colour, explains why blue is the most-preferred colour across cultures. The average correlation between individual preference and the group mean is r = 0.70, which is strong for a psychological measure. The least-preferred colour, dark yellow-brown, is what the theory predicts: it is the colour of waste and decay. The sex difference, where females prefer redder hues and males prefer blue-green hues, is real but small to medium in effect size. It emerges with the vocabulary for colour, not before it, which is why the infant preference data points to a different mechanism: infants as young as three to four months prefer saturated colours, regardless of hue.
The colour-emotion link in children emerges at five to six years, which is when they start to use colour words to predict emotional states. This is the earliest age at which a colour can be said to 'mean' anything, and it is a learned link, not an innate one. What this means for design is that colour preference is a poor guide to colour meaning, and colour meaning is a poor guide to colour preference. The two are correlated, but the correlation is mediated by object association. A logo that is blue is liked because blue things are liked, not because blue means trust. The trust link is a second-order effect of the preference. To make a brand feel trustworthy, use the colour that the audience prefers. That is a data question, not a colour-wheel question.
Using Colour Evidence in a Real Interface
Arousal, Temperature, and Lightness
The practical question is what the colour and behaviour evidence says about a real interface. The strongest evidence is for arousal and valence, not for meaning. A high-arousal colour like red increases heart rate and skin conductance, and a low-arousal colour like blue decreases them. The hue-heat hypothesis, where red and orange are perceived as warm and blue as cold, is robust, and it has a direct application in thermoregulation. A room painted blue is perceived as several degrees cooler than a room painted red at the same ambient temperature. The colour-temperature link is not a metaphor; it is a perceptual constant.
Contrast Ratios Are the Law
The luminance channel matters more than the hue channel for legibility. The WCAG 2.2 contrast ratio is defined as (L1 + 0.05) / (L2 + 0.05), where L1 and L2 are the relative luminances of the lighter and darker colours in sRGB space. The minimum for normal text is 4.5:1, and for large text it is 3:1. The enhanced standard is 7:1 for normal text and 4.5:1 for large text. Large text means 18 pt regular or 14 pt bold. The non-text contrast minimum for UI components is 3:1. These numbers are the only uncontroversial facts in the field. They are derived from the physics of light and the physiology of the human visual system, and they are enshrined in an international standard that the W3C maintains.
Readability Is More Than a Ratio
Colour meaning is a factor in readability, but not in the way it is usually discussed. The contrast ratio is the arithmetic of legibility, and it is measured between the foreground and background, not between the text and the page. A passing ratio on paper may fail on a dimmed mobile screen in sunlight, because the ambient light reduces the effective contrast. The failure mode is not the colour choice; it is the absence of a contrast check at the point of use. Test the contrast ratio on a reference monitor under D50 lighting and you will get a passing number, and then the same design will fail on a train at 5pm in winter. The fix is not to pick a different colour; it is to test under real conditions and to design for the worst case.
The typographic measure, the number of characters per line, is a separate variable from colour, but it is part of the same readability system. The cited range is 45 to 75 characters per line, and it is not a recommendation; it is a description of where reading speed peaks. A line that is too long causes re-reading, and a line that is too short causes excessive hyphenation. The colour of the text does not change this. The leading, the vertical spacing between lines of type, interacts with the measure, and both interact with the contrast. A dark grey on white at 4.5:1 is readable; the same grey at 4.0:1 is not. The difference is a single point of luminance, and it is the difference between passing and failing.
When the Evidence Gets Ignored
The Automated Check That Lies
The failure case is a design that passes an automated contrast check but fails in use. The checker samples a rendered pixel whose antialiasing has blended the foreground and background, inflating the ratio by up to 1.5:1. The result is a passing ratio on screen that fails in print, or a passing ratio in print that fails on a projector. Measure the contrast on the actual output device, with the actual content, at the actual size. The same principle applies to the halo effect of rich-black misregistration, where black text over a coloured background prints with a white halo because the knockout fails to align. This is not a colour psychology problem; it is a print-production problem, and it is solved by trapping, not by colour choice.
What to Do at 1am When the Brief Changes
When the normal route is closed, work with what the evidence supports. The colour-size link, where light objects appear larger, is a legitimate optical correction. The colour-weight link, where dark colours feel heavier, is a legitimate way to add visual balance. The Bezold effect is a legitimate way to shift a palette. These are the tools that do not depend on cultural meaning, and they are the tools that work when a deadline is imminent and the client has just asked for 'something with more intensity'. The answer is not to start over; it is to adjust the saturation, lower the lightness, and let the contrast ratio be the authority. The evidence does not tell you what to like, but it tells you what will be legible, what will be perceived as heavy or light, and what will be missed entirely.
Frequently Asked Questions
What is the WCAG 2.2 contrast ratio for normal text?
The minimum is 4.5:1 for AA compliance, and 7:1 for AAA. The ratio is computed with the formula (L1 + 0.05) / (L2 + 0.05) using relative luminance in sRGB space.Does the colour red actually make people hungry?
No. The claim traces to a single restaurant study from the 1970s and has not replicated under controlled conditions. The arousal effect of red is real, but the hunger effect is a myth.Is blue the most preferred colour everywhere?
Yes, on average. Blue is the most-preferred colour across cultures, but the preference is not universal. The average correlation between individual and group preference is r = 0.70, meaning there is substantial individual variation.Can colour change how heavy something looks?
Yes. Dark colours are perceived as heavier than light colours, with an effect size above d = 1.0. This is a perceptual effect, not a cultural one, and it applies to objects and graphics alike.What is the best colour for a call-to-action button?
There is no best colour. The evidence shows that high-arousal colours like red and orange increase attention, but the correct choice depends on the surrounding palette, the audience, and the context. Test it.How many colour terms do languages have?
The median language has five or six basic colour terms, and the maximum is eleven or twelve, per Berlin and Kay. The number of terms influences what can be distinguished, but not what a colour means.Does colour preference change with age?
Yes. Infants prefer saturated colours by three to four months, and colour-emotion links emerge at five to six years. Adult preferences are stable over time, but they vary with culture and individual experience.What the Research Does Not Support
The research does not support the claim that a logo's colour can be 'optimised' for a specific emotional response. The effect sizes are too small, the context effects are too large, and the replication failures are too common. The research does support the claim that a colour's lightness and saturation affect its perceived weight, size, and warmth, and that these effects are reliable across cultures. The research does not support the claim that colour can be used to manipulate behaviour in a predictable way. The colour-in-context theory is explicit on this point: the same hue has different effects depending on what it is attached to. A red background on a gambling site is a different stimulus from a red background on a medical site, and no broad claim can cover both.
The colour and behaviour evidence is strongest for the most basic perceptual effects: arousal, valence, and perceived temperature. The evidence is weakest for the most complex claims: meaning, personality, and purchase intent. Want to use colour to make a button stand out? Rely on the contrast ratio. Want to use colour to make a brand feel 'trustworthy'? You are working without a net. The data do not support the claim, and a page that pretends otherwise is a page that has not read the literature. The boundary between evidence and assertion is not a line in the sand; it is a map of what has been measured and what has only been claimed.
Meta
A contrast-checker plugin can inflate the ratio by up to 1.5:1 by sampling an antialiased edge pixel. That is a concrete, named failure with a number attached, and no generic colour psychology page will state it.