A Likert scale looks trivial and is easy to get wrong. Unbalanced anchors, a missing midpoint, or uneven spacing between options all bias responses, and none of it can be repaired after data collection.
Choose what you are measuring and how many points you want, and this tool produces a properly balanced set of anchors along with the scoring rules, the analysis that scale supports, and the item-writing rules that decide whether your data is worth analysing.
A balanced scale has the same number of positive and negative options, and the wording is symmetrical: if one end is "strongly agree", the other must be "strongly disagree", not merely "disagree". Anchors like Excellent, Very good, Good, Fair place three favourable options against one unfavourable and will push your mean upward regardless of what respondents think.
Spacing matters too. Respondents read the options as roughly evenly spaced, which is what licenses treating a summed scale as interval data. Anchors that jump from "never" to "always" with "sometimes" as the only middle break that assumption.
Five points is the default because it offers a midpoint and enough range without overtaxing the respondent. Seven gives finer discrimination and slightly more reliable summed scores, at the cost of a longer form and more effort per item — reasonable for an established instrument, less so for a first questionnaire.
Four points removes the midpoint and forces a direction. That is a legitimate choice when you believe respondents use the midpoint to avoid deciding, but it is a real cost: someone with genuinely no view has nowhere honest to go, and they will either guess or abandon the survey. Whichever you choose, justify it in your methodology and keep it consistent across the whole instrument.
Code the options 1 to n in order. Where an item is worded negatively — included deliberately to catch respondents ticking straight down one column — reverse its codes before analysis, or it will pull against the items it is meant to reinforce. Forgetting to reverse-code is one of the most common reasons a scale shows poor reliability.
Where several items measure one construct, report Cronbach's alpha for the set. Above 0.70 is the usual threshold for internal consistency. A low alpha usually means the items are not measuring the same thing, or that a reverse-coded item was missed.
A single Likert item is ordinal. The gap between "agree" and "strongly agree" is not demonstrably the same as the gap between "neutral" and "agree", so the mean of one item is hard to defend. Report the median and the frequency of each option instead.
A Likert scale — several items summed or averaged to measure one construct — is conventionally treated as interval, and that is what permits t-tests, ANOVA, correlation and regression. The distinction between a Likert item and a Likert scale is worth stating explicitly in your methodology, because examiners do ask about it.
One idea per item. "The service is fast and affordable" cannot be answered by someone who finds it fast and expensive, and you will never know which half they responded to. Avoid leading wording, double negatives, and technical vocabulary your respondents may not share.
Pilot the wording with five or six people from your target population and watch for hesitation or re-reading — both signal ambiguity. Fixing an item after a pilot costs an afternoon; discovering the problem in your results chapter costs a great deal more.
Five is the common default and easier to complete. Seven offers finer discrimination and slightly better reliability for summed scales, but makes the form longer.
Usually yes — without one, respondents with no genuine view are forced to take a side. Removing it is defensible if you can justify it, but say so in your methodology.
For a single item, prefer the median: one item is ordinal. For several items summed into a scale measuring one construct, a mean is conventional and supports parametric tests.
Flipping the scores of negatively worded items so they point the same way as the rest. Omitting this step is a frequent cause of unexpectedly low reliability.
Typically four to eight. Fewer makes reliability hard to demonstrate; many more tires respondents without adding much.