TL;DR — Quick Answer
A Likert scale is a rating scale used to measure attitudes, opinions, and perceptions by asking respondents how strongly they agree or disagree with a series of statements — most commonly on a 5-point scale from “Strongly disagree” to “Strongly agree.” Strictly, a single statement is a Likert item; the Likert scale is the combined score across several related items measuring the same construct. Likert data is ordinal at item level, which shapes how it should be analysed: use medians, frequencies, and bar charts for individual items, while summed multi-item scales are widely treated as approximately interval data suitable for means, correlations, and parametric tests.
Open almost any survey-based thesis or journal article in management, education, psychology, or the social sciences, and you will find a Likert scale doing the measurement work. It is the most widely used attitude-measurement instrument in research — and also one of the most frequently misused. This guide explains what a Likert scale actually is, how to design one properly, how to analyse the data it produces, and the debates every researcher using one should understand before facing an examiner or reviewer.
What Is a Likert Scale?
The Likert scale was developed by the American social psychologist Rensis Likert in 1932 as a technique for measuring attitudes. Its logic is simple: instead of asking people directly to quantify an attitude (“How satisfied are you, on a scale of your choosing?”), it presents a series of clear statements and asks respondents to indicate their level of agreement with each, on a symmetric scale with a defined number of points.
The classic 5-point agreement format is:
- 1 — Strongly disagree
- 2 — Disagree
- 3 — Neither agree nor disagree
- 4 — Agree
- 5 — Strongly agree
The same structure adapts to other dimensions: frequency (Never → Always), satisfaction (Very dissatisfied → Very satisfied), importance (Not at all important → Extremely important), and likelihood (Very unlikely → Very likely). What defines the format is not the word “agree” but the symmetric, ordered, labelled response continuum.
Likert Item vs Likert Scale — A Distinction Examiners Notice
In careful usage, a single statement with its response options is a Likert item. A Likert scale, strictly speaking, is the composite score produced by summing or averaging a respondent’s answers across several related items that together measure one underlying construct.
The distinction matters practically. Suppose you are measuring “job satisfaction.” One item — “I am satisfied with my job” — captures a narrow slice and is vulnerable to momentary mood and interpretation. A scale of six items covering pay, workload, recognition, colleagues, growth, and overall contentment produces a more stable, more valid measure of the construct. Most established research instruments — the kind you will encounter when reviewing published questionnaires — are multi-item scales for exactly this reason, and the internal consistency of those items is what reliability statistics assess. This is also why single-item measures attract methodological criticism in viva examinations and peer review: they measure less reliably, and there is no way to check their consistency internally.
How Many Points Should the Scale Have?
The 5-point and 7-point formats dominate published research, and the evidence on which is “better” is genuinely mixed — meaning your choice should be justified, not defaulted.
- 5-point scales are easier for respondents, translate cleanly across languages, and reduce completion time — meaningful advantages for long questionnaires and general populations.
- 7-point scales offer finer discrimination and slightly better distributional properties for statistical analysis, and suit educated or professional respondent groups comfortable with nuance.
- Even-numbered scales (4 or 6 points) remove the neutral midpoint, forcing a directional response. This is a deliberate design decision: it prevents fence-sitting but frustrates respondents who genuinely hold no position, and can distort data by manufacturing opinions that do not exist.
- Beyond 7 points, gains in precision largely vanish — respondents cannot reliably distinguish point 8 from point 9 on an 11-point continuum.
Whichever you choose, two rules are non-negotiable: keep the number of points consistent across the questionnaire (switching between 5- and 7-point sections invites response errors and complicates analysis), and label the points clearly — at minimum the endpoints and midpoint.
Writing Good Likert Items
The scale is only as good as its statements. The craft principles overlap with questionnaire design generally — covered in depth in our guide on how to design a research questionnaire — but the ones specific to Likert items are worth stating:
- One idea per statement. “My supervisor is supportive and competent” is a double-barrelled item — a respondent may agree with one half and not the other. Split it.
- Avoid absolutes and negations. “I never feel stressed at work” combines an absolute (“never”) with a construct better measured positively; and double negatives (“I do not disagree that…”) are uninterpretable.
- Keep statements attitudinal, not factual. “The office opens at 9 a.m.” has a correct answer; agreement scales measure positions, not facts.
- Use reverse-coded items deliberately and sparingly. Including some negatively phrased items (“I often think about leaving this job” within a satisfaction scale) helps detect inattentive straight-line responding — but they must be re-coded before analysis, and poorly translated reverse items are a notorious source of reliability problems.
- Match the response anchors to the statement. Frequency statements need frequency anchors, not agreement anchors. “I check email after work hours — Strongly agree” is a mismatch; “— Always/Often/Sometimes/Rarely/Never” is correct.
Is Likert Data Ordinal or Interval? The Debate You Must Be Able to Discuss
This is the methodological question examiners love, and it deserves an honest treatment rather than a dodge.
The strict position: Likert responses are ordinal. The categories are ordered, but the psychological distance between “Strongly agree” and “Agree” is not demonstrably equal to the distance between “Agree” and “Neutral.” Calculating a mean of ordinal codes therefore rests on an assumption the data cannot verify, and the technically appropriate statistics are medians, frequencies, and non-parametric tests.
The pragmatic position: when several Likert items are summed or averaged into a composite scale score, the resulting variable takes many values, tends toward a roughly continuous distribution, and behaves — for practical analytical purposes — like interval data. Decades of methodological work have shown that parametric tests (t-tests, ANOVA, correlation, regression) are robust when applied to multi-item Likert scale scores, and the overwhelming majority of published survey research proceeds on this basis.
The defensible synthesis, and the one worth writing into a methodology chapter: treat individual items as ordinal; treat validated multi-item composite scores as approximately interval. Report item-level results with medians and frequency distributions; run parametric analyses on composite scores, checking distributional assumptions as you would for any variable. Whichever position you adopt, state it explicitly and justify it — the error is not choosing a side, it is appearing unaware that a question exists. For the underlying concepts, see our guides on variables in research and descriptive vs inferential statistics.
How to Analyse Likert Data
A sensible analysis sequence for survey data built on Likert measurement:
- Step 1 — Screen and re-code. Check for missing responses and straight-lining; reverse-code negatively worded items so that higher values consistently mean “more” of the construct.
- Step 2 — Assess reliability. For each multi-item scale, compute an internal-consistency statistic (Cronbach’s alpha is the standard) and report it; values of 0.70 and above are conventionally acceptable for established scales.
- Step 3 — Describe item-level results. Report frequencies or percentages per response category, with the median or mode as the summary statistic. Diverging stacked bar charts are the clearest visualisation for agreement items; pie charts are poorly suited to ordered categories.
- Step 4 — Compute composite scores. Sum or average each respondent’s items per scale, and describe the composites with means and standard deviations.
- Step 5 — Test your hypotheses on the composites. Group comparisons (t-tests, ANOVA or their non-parametric counterparts Mann–Whitney and Kruskal–Wallis), correlations, and regression models all operate on the composite scores. Choosing among these depends on your research questions and data properties — our guide on how to choose the right statistical test maps the decision.
Common Mistakes to Avoid
- Averaging a single item and reporting it to two decimal places — false precision on ordinal data; use the median or the percentage distribution instead.
- Mixing scale lengths mid-questionnaire without analytical justification.
- Forgetting to reverse-code negatively worded items — the single most common cause of mysteriously low reliability coefficients.
- Treating the neutral midpoint as “no data.” A genuine “neither agree nor disagree” is information; excluding midpoint responses biases results.
- Borrowing a published instrument and modifying items silently. Any alteration — wording, anchors, number of points — means the original validation no longer fully applies, and this must be acknowledged and ideally re-tested.
Frequently Asked Questions
Is a Likert scale qualitative or quantitative?
Quantitative. It converts subjective attitudes into ordered numerical data for statistical analysis. The construct being measured is subjective; the data produced is numerical.
Can I use a mean with Likert data?
For composite multi-item scale scores, yes — this is standard practice. For a single item, the median and frequency distribution are the more defensible summaries, though means on single items remain common in published work; if you report one, do so alongside the distribution.
What is the difference between a Likert scale and a rating scale?
Every Likert scale is a rating scale, but not every rating scale is a Likert scale. The Likert format specifically uses symmetric, labelled agreement-style categories in response to statements; a 1–10 “rate your experience” slider or a semantic differential (bipolar adjectives such as “Modern … Traditional”) are different rating formats with their own properties.
How many items should a Likert scale contain?
Enough to cover the construct’s facets while respecting respondent fatigue — most validated scales run between 4 and 10 items per construct. Fewer than three items makes internal-consistency assessment weak; beyond ten, marginal items usually add fatigue faster than information.
Final Thoughts
The Likert scale endures because it solves a genuinely hard problem — turning private attitudes into analysable data — with an instrument respondents understand in seconds. Used well, it is the backbone of credible survey research; used carelessly, it produces confident-looking numbers built on double-barrelled items, unjustified means, and unexamined assumptions. Design items that isolate one idea each, keep the format consistent, know where you stand on the ordinal–interval question, and analyse items and composites at their appropriate levels. Do that, and the most familiar instrument in social research will serve you exactly as Rensis Likert intended.