What is the scientific method in psychology? It is the way psychologists test ideas about behaviour and mental processes: they turn a question into a testable hypothesis, measure the variables with clear definitions, collect evidence from a sample of people, analyse the results with statistics and share them so others can check. Because people are variable and cannot be fully controlled, psychology adds extra safeguards that a physics experiment rarely needs.
Why psychology needs the scientific method
Everyone has theories about people. "Teenagers need more sleep." "Revision works better with music." "Eyewitnesses always remember what they saw." Some of these ideas are right, some are wrong, and a few are only partly true. Common sense cannot tell them apart, because we notice the cases that fit our beliefs and forget the ones that do not.
Psychology is the scientific study of behaviour and mental processes, and its claim to be a science rests on its methods rather than its subject. Psychologists make claims that can be tested, measure what they claim to study, and accept that evidence can prove them wrong. A claim that no possible result could ever contradict is not scientific. Philosopher Karl Popper called the ability to be proved wrong falsifiability.
This article is about what makes the method distinctive in psychology: hypotheses, operational definitions, research designs, sampling, ethics and the replication crisis. If you want the general step-by-step routine used across all sciences, read our guide to how to use the scientific method first. Here we stay with the problems that only arise when the thing you are studying is a person.
From research question to hypothesis
Psychological research starts with an observation or a puzzle. A good research question is specific, measurable and testable. "Does sleep affect memory?" is a topic. "Do students who sleep for 8 hours recall more words than students who sleep for 5 hours?" is something you can actually investigate.

The infographic above shows the two steps side by side. The next job is to turn the question into a hypothesis: a precise, testable prediction about what will happen. A hypothesis is not a wild guess. It normally comes from existing theory or earlier research, which is why psychologists read the literature before they collect any data.
Directional, non-directional and null hypotheses
Psychology courses, including GCSE and A-level, expect you to tell three kinds of hypothesis apart.
- Directional (one-tailed): predicts which way the difference or relationship goes. "Students who sleep for 8 hours will recall more words than students who sleep for 5 hours."
- Non-directional (two-tailed): predicts a difference but not its direction. "There will be a difference in the number of words recalled by students who sleep for 8 hours and students who sleep for 5 hours."
- Null: predicts no difference or relationship, apart from chance. "There will be no difference in words recalled between the two sleep groups."
Researchers choose a directional hypothesis when earlier findings point a clear way, and a non-directional one when the evidence is mixed. Statistical tests then ask how likely the observed results would be if the null hypothesis were true. If that chance is small enough, the null hypothesis is rejected.
Variables and operational definitions
A variable is anything that can change or be measured. In an experiment the researcher changes the independent variable (IV) and measures the dependent variable (DV). In our example the IV is the amount of sleep and the DV is the number of words recalled.
The trouble is that many psychological ideas, such as stress, intelligence, happiness or memory, cannot be seen or weighed. To study them, researchers write an operational definition: a precise statement of exactly how the variable will be manipulated or measured. Without one, two studies can use the same word and mean different things.
| Concept | Vague idea | Possible operational definition |
|---|---|---|
| Memory | "How well people remember" | Number of words correctly recalled from a list of 20 after a 10-minute delay |
| Sleep | "A good night's rest" | Time in bed, recorded as 8 hours or 5 hours, checked by a sleep diary |
| Stress | "Feeling under pressure" | Score on a standard stress questionnaire, or a measured stress hormone level |
| Aggression | "Acting angry" | Number of times a child hits a doll in a 20-minute observation, counted by trained observers |
Operational definitions make a study replicable: another team can follow the same recipe. They also have a cost. Reducing "memory" to a word list captures only one slice of what memory is, so a good researcher admits this limit in the report.

Extraneous and confounding variables
Anything other than the IV that could affect the DV is an extraneous variable. If it actually varies along with the IV, so that you cannot separate the two effects, it becomes a confounding variable. Suppose the 8-hour group were tested in the morning and the 5-hour group in the afternoon. Any difference in recall could come from sleep or from time of day, and the study could not tell which.
Researchers deal with these variables by standardising procedures (same instructions, same room, same time of day), by using control groups and by allocating people to conditions at random. We cover these tools in the next part. For a look at how sleep loss really affects the body and mind, see our article on sleep deprivation effects.
Why a "good hypothesis" must be falsifiable
A claim such as "everyone has hidden desires they are not aware of" sounds deep, but it can never be tested: if a person shows no sign of a desire, the claim just says it is hidden. Scientific hypotheses stick their necks out. "People who sleep for 5 hours will recall fewer words than people who sleep for 8" can clearly fail. If the groups score the same, the hypothesis loses.
This is also why psychologists do not talk about proving a hypothesis. A study can support it. A different sample, a different measure or a different country may later give a different answer. Science builds confidence step by step, which brings us to how studies are designed.
Experimental designs: testing cause and effect
Only one kind of study can show that one variable causes a change in another: the experiment. The researcher manipulates the IV, measures the DV and keeps everything else as constant as possible. Participants are usually split into an experimental group, which receives the treatment, and a control group, which does not, so the two can be compared.
Types of experiment
- Laboratory experiment: a controlled setting where extraneous variables are tightly managed. High control, but behaviour may be less natural.
- Field experiment: the IV is manipulated in a real-life setting, such as a school or station. More natural behaviour, less control.
- Natural experiment: the IV changes by itself, for example before and after a new law, and the researcher simply measures the effect.
- Quasi-experiment: the IV is a feature of the participants, such as age or gender, so the researcher cannot assign people to it.
Three ways to use participants
| Design | What happens | Strength | Weakness |
|---|---|---|---|
| Independent groups | Different people in each condition | No practice or fatigue effects | Differences between people may blur the result |
| Repeated measures | The same people do every condition | Removes individual differences | Order effects such as practice and boredom |
| Matched pairs | Pairs matched on a key trait, one in each condition | Controls the matched trait | Time-consuming; matching is never perfect |
In repeated-measures studies, researchers use counterbalancing: half the people do condition A first and half do B first, so order effects cancel out. In independent-groups studies they use random allocation, for instance drawing names from a hat, so that each person has the same chance of ending up in either group.
Correlational studies: patterns without proof of cause
Often researchers cannot or should not manipulate a variable. You cannot assign children to be bullied or adults to have a stressful life. Instead they measure two variables as they naturally occur and ask whether they are linked. This is a correlational design.
The strength and direction of the link is given by a correlation coefficient, a number between −1 and +1. A value near +1 means that as one variable rises the other rises. A value near −1 means one rises as the other falls. A value near 0 means there is no linear relationship. A scatter graph shows the pattern.
| Feature | Experiment | Correlational study |
|---|---|---|
| Researcher manipulates a variable? | Yes (the IV) | No |
| Can show cause and effect? | Yes, if well controlled | No |
| Typical result | Difference between conditions | Correlation coefficient |
| Example | Sleep 5 h or 8 h, then a memory test | Average hours of sleep compared with exam marks |

Other research methods
As the infographic shows, psychologists also use methods that describe behaviour rather than test cause and effect.
- Observation records what people do, either naturally (a playground) or in a controlled setting. Observers use behavioural categories and tally charts so that two observers record the same thing. Agreement between them is called inter-observer reliability.
- Surveys and questionnaires collect responses from many people quickly. Wording matters, and people sometimes answer in the way they think looks good, which is called social desirability bias.
- Interviews allow deeper answers but take longer, and the interviewer can influence the replies.
- Case studies examine one person or small group in depth, often over time. They are rich in detail but hard to generalise from.
Strong research often combines methods. Brain-imaging studies, for instance, add physical evidence to behavioural data. Our article on neuroscience and the brain explains how scans link to behaviour.

Making sense of the data
Once data are collected, researchers summarise them with descriptive statistics: measures of central tendency (mean, median, mode) and measures of spread (range, standard deviation). Tables and graphs show patterns. Inferential statistics then test whether a difference or relationship is bigger than chance would predict. By convention, psychologists usually accept results as significant when the probability of getting results at least this extreme, if there were really no effect (the null hypothesis), is below 5% (p < 0.05).
Try it with five recall scores from a small sleep study. Change the numbers and see how the mean and range respond. A value far from the rest is flagged as a possible anomaly: check why it happened before you remove it.
Significance has a catch. "Statistically significant" does not mean "large" or "important". With a huge sample, a trivial difference can be significant, so researchers also report the effect size. We return to this when we reach the replication crisis.
Sampling: who takes part?
No study can test everyone. The target population is the group the researcher wants to draw conclusions about, such as all UK teenagers. The sample is the smaller group who actually take part. If the sample does not resemble the population, the findings may not generalise.
| Sampling method | How it works | Main problem |
|---|---|---|
| Random | Every member of the population has an equal chance of selection | Needs a full list of the population; may still miss groups by chance |
| Stratified | Sample mirrors the population's subgroups (for example age bands) | Time-consuming to organise |
| Opportunity | Whoever is available and willing at the time | Often biased, for example only one school |
| Volunteer (self-selected) | People respond to an advert | Volunteers tend to be more motivated than average |
Much classic research used university students. In 2010 Joseph Henrich, Steven Heine and Ara Norenzayan pointed out in Behavioral and Brain Sciences that most behavioural science samples come from people in Western, educated, industrialised, rich and democratic societies. They called such samples WEIRD and argued that findings from them may not apply to everyone.
Sample size matters too. Very small samples can produce striking results by luck, and they have little statistical power to detect real effects. For practical, study-focused advice on the topic our sleep example tests, see our guide to how to improve memory for better learning.

Classic psychology studies and what they tested
Four famous studies show the method at work. Each one used a different design, and each raises a point about evidence or ethics. The figures below come from the original reports as summarised in standard psychology references.
Loftus and Palmer (1974): can wording change memory?
Elizabeth Loftus and John Palmer asked whether the wording of a question could change what people remember about an event. This was a laboratory experiment with an independent-groups design. In the first experiment, 45 students watched seven short films of traffic accidents and then estimated the cars' speed. The IV was the verb in the question ("About how fast were the cars going when they ___ each other?"), and the DV was the speed estimate in miles per hour.
The verb changed the answers. Groups asked about cars that smashed estimated 40.8 mph on average, while groups asked about cars that contacted gave 31.8 mph. In a second experiment with 150 students, people were asked a week later whether they had seen broken glass, and there was none in the film. Of those who had heard "smashed", 32% said yes, compared with 14% for "hit" and 12% for the control group.
The study does not show that all eyewitness testimony is worthless, which is a claim often repeated. It shows that leading questions can alter reports of what was seen. Its weakness is ecological validity: watching a film in a lecture room is not the same as witnessing a real crash.
Asch (1951): do groups pressure us?
Solomon Asch tested conformity with a simple task: choose which of three lines matched a standard line. In his best-known set-up each participant sat with a group in which everyone else was a confederate, an actor working for the researcher. Over 18 trials, the confederates gave the wrong answer on 12 critical trials. About 74% of participants went along with the group at least once, and about 12% did so on nearly every critical trial. In a control condition with no group pressure, participants made errors in less than 1% of trials.
The task had an obvious correct answer, so Asch could measure conformity cleanly. That is a strength of the design, but it is also its limit: the lines mattered little to anyone, and the original groups were made up of male students. For everyday pressure from friends, see negative and positive peer pressure with examples.
Bandura (1961): is aggression learned by watching?
Albert Bandura's team at Stanford studied 72 children aged 3 to 6 from a university nursery school. Children watched an adult model behave aggressively toward an inflatable Bobo doll, watched a non-aggressive model, or saw no model (the control group). Afterwards each child was observed alone with toys. Children who had watched the aggressive adult copied both physical and verbal aggression more than the others.
This was a laboratory experiment with observation as the measurement. It supported the idea that behaviour can be learned by imitation without any reward. It also raised ethical questions, which we meet shortly: children were deliberately exposed to aggression and could not give their own informed consent.
Milgram (1963): how far does obedience go?
At Yale University, Stanley Milgram told 40 men aged 20 to 50 that they were helping with a study of memory and learning. They were told to give a "learner" an electric shock whenever he made a mistake, with the shock level rising to 450 volts. The learner was an actor and no shocks were delivered. About 65% of participants went all the way to 450 volts, although many showed obvious distress.
The study is a landmark in social psychology and a landmark in research ethics. Participants were deceived about the true purpose, were exposed to severe stress, and were encouraged to continue when they hesitated. Milgram did debrief everyone afterwards and later surveyed them, and most said they were glad they had taken part. Nevertheless, studies like this one would not be approved today in the same form.
Ethics in psychological research
Some of the most important rules in psychology were written because of studies like those above. In the UK, researchers follow the British Psychological Society (BPS) Code of Human Research Ethics, whose current edition was published in 2021. It is built on principles that include respect for the autonomy and dignity of persons, scientific integrity, social responsibility, and maximising benefit while minimising harm. University and school research must be approved by an ethics committee before any participant is recruited.
Informed consent, deception and children
Informed consent is the heart of the code, yet it can clash with the needs of the study. If a participant knows that Asch's other group members are actors, the experiment is pointless. Researchers therefore sometimes use mild deception, but only when no alternative exists, when the risk is low, and when a full debrief follows. Some studies ask for consent in advance for taking part in "a study that may involve some deception" without revealing the details.
When the participants are children, consent for younger children (usually under 16) normally has to come from a parent or carer, as well as the child's own agreement, and the child's own willingness to take part is respected throughout. Participants who cannot give consent for themselves, or who are particularly vulnerable, need extra safeguards.
Ethics for students
If you run a practical for GCSE or A-level, the same ideas apply on a smaller scale. Tell people what the study involves, let them stop at any moment, do not collect names unless you must, and explain the purpose afterwards. Our overview of ethical dilemmas in science and technology places these questions in a wider setting.
- Participants know what the task involves before they start
- Consent is given freely, with parent or carer consent for children
- Participants know they can stop at any time
- No pressure, distress or embarrassment beyond everyday life
- Names and personal details are not recorded unless essential
- A debrief explains the real aim and any deception
- Teacher or ethics committee has approved the plan
The replication crisis: can psychology's findings be repeated?
A finding is only trustworthy if other researchers can repeat the study and get a similar result. That is replication, the final step in the infographic below. In the 2010s psychologists discovered that many well-known results were harder to repeat than expected, which became known as the replication crisis.

The Open Science Collaboration (2015)
The landmark test was led by Brian Nosek and a large team called the Open Science Collaboration. They attempted to repeat 100 studies published in 2008 in three leading journals: Psychological Science, the Journal of Personality and Social Psychology and the Journal of Experimental Psychology: Learning, Memory, and Cognition. Their report appeared in Science in 2015.
- 97 of the 100 original studies reported a significant result.
- Only 36% of the replications produced a significant result.
- On average, the effects found in the replications were about half the size of the original effects.
It is easy to over-read these numbers. They do not prove that 64% of the original findings were false. A replication can fail because of a different sample, small differences in procedure, or plain chance. The authors themselves said that neither a single success nor a single failure settles an effect. What the project showed is that published significant results were less reliable than most people assumed, and that the way psychology was practised needed to change.
Why did so many results fail to replicate?
- Publication bias: journals prefer exciting, significant findings, so null results often stay in a drawer.
- Low statistical power: small samples can produce large effects that vanish with more participants.
- Questionable research practices: trying many analyses and reporting only the one that works (p-hacking), or deciding the hypothesis after seeing the data.
- Career pressure: "publish or perish" rewards new, positive findings more than careful checks.
A famous example of a surprising claim under scrutiny was Daryl Bem's 2011 paper reporting evidence for precognition, the ability to sense future events. Other researchers, including Stuart Ritchie, Richard Wiseman and Christopher French, tried to repeat one of the experiments and did not find the effect. The episode made many psychologists ask how a study that followed accepted methods could produce such an unlikely result, which led them to examine their own standards.
How psychology is fixing itself: pre-registration and open science
Pre-registration means writing down your hypothesis, design, sample size and planned analysis in a public registry before collecting data. The Center for Open Science explains that the same data cannot be used both to generate and to test a hypothesis. Pre-registration separates planned (confirmatory) tests from exploratory ones, so readers know which results to trust most.
A step further is the Registered Report. Here a journal peer-reviews the question and method before the results exist. If the plan is sound, the journal gives "in-principle acceptance" and agrees to publish the paper whatever the outcome, which removes the pressure to chase significant results.
Other reforms
- Open data and materials: sharing questionnaires, code and data so others can check them.
- Larger samples: planning sample size in advance for adequate power.
- Replication studies: treating repeats as valuable work, not as copies with no credit.
- Reporting effect sizes as well as p-values, so readers can judge how big an effect is.
None of this means psychology is not a science. A field that finds its own weaknesses with its own methods and changes its practice is behaving the way a science should. It does mean that one new study is a clue, not a verdict.
How to judge a psychology claim
You will meet psychology claims in headlines, social media and textbooks. A few questions help you decide how much to trust them. They also make good evaluation points in an exam.
- Was it an experiment (can show cause) or a correlation (cannot)?
- How many people took part, and who were they?
- Were the key ideas given operational definitions?
- Has anyone replicated it, ideally with a different sample?
- Was it pre-registered, and are the data shared?
- Was it conducted ethically?
For findings that stand up to such checks, see our overview of positive psychology facts, and for the general version of the same logic, read scientific method made easy.
Test yourself
Explain why a correlational study cannot show that one variable causes another. (3 marks)
The researcher does not manipulate any variable, so there is no control over other factors. A third variable could explain the link. Even if the link is real, the correlation does not show which variable influences the other.
Write an operational definition of "stress" for a student project. (2 marks)
For example: "Stress is the score on a ten-item stress questionnaire completed by each participant, where higher scores show more stress." It states exactly how the variable is measured.
Give two ethical issues in Milgram's obedience study. (2 marks)
Participants were deceived about the true purpose and the learner's shocks, and they experienced considerable stress. Their right to withdraw was undermined because they were told to continue when they hesitated.
Frequently asked questions about the scientific method in psychology
What is the scientific method in psychology?
It is the systematic way psychologists test ideas about behaviour and mental processes. They form a testable hypothesis, define and measure variables, collect data from a sample, analyse it statistically, draw cautious conclusions and share their methods so that others can replicate the study.
What are the steps of the scientific method in psychology?
Typical steps are observation, a research question, a hypothesis, designing and running the study, analysing the results, and reporting and replicating. In psychology, extra attention goes to operational definitions, sampling and ethical approval before any participant takes part.
What is an operational definition in psychology?
An operational definition states exactly how a variable is measured or manipulated. For example, "memory" might be defined as the number of words correctly recalled from a list of 20 after ten minutes. It allows other researchers to repeat the study in the same way.
What is the difference between an experiment and a correlational study?
In an experiment the researcher manipulates the independent variable and measures its effect, so cause and effect can be tested. In a correlational study two variables are measured as they naturally occur, so the researcher can describe a link but cannot say that one causes the other.
Why is a hypothesis needed in psychology?
A hypothesis turns a general question into a precise prediction that can be supported or contradicted by evidence. It guides the design, tells the researcher what to measure and prevents them from deciding what the results mean only after seeing the data.
What is the replication crisis in psychology?
It is the discovery, mainly in the 2010s, that many published psychology findings are hard to repeat. In the 2015 Open Science Collaboration project, only 36% of 100 replications gave significant results, compared with 97 of 100 original studies, and effects were about half the original size.
What is pre-registration in psychology?
Pre-registration means publicly recording your hypothesis, method and analysis plan before collecting data. It separates planned tests from exploratory ones and makes it harder to change the analysis after seeing the results. Registered Reports go further by having journals review the plan first.
What is the BPS Code of Human Research Ethics?
It is the British Psychological Society's guidance for ethical research with human participants, most recently published in 2021. It covers respect for autonomy and dignity, scientific integrity, social responsibility, and maximising benefit while minimising harm, including informed consent, the right to withdraw and debriefing.
Why is informed consent important in psychology research?
Informed consent lets people decide freely whether to take part after learning what is involved. It respects their autonomy and helps protect them from harm. If deception is unavoidable, researchers must justify it, keep risks low and fully debrief participants afterwards.
Can psychology be a real science if results do not always replicate?
Yes. Science is judged by its methods and its willingness to correct itself, not by getting every result right first time. The replication crisis was found by psychologists using scientific methods, and reforms such as pre-registration and open data came from within the field.



