Entry
Reader's guide
Entries A-Z
Subject index
Reliability
Introduction
Reliability as a central concept of test theory dates back to the beginning of the 20th century. It is based on the existence of intra-individual variability as well as variation between persons. With intra-individual variability or measurement error, true score was also introduced as a central concept of classical test theory. Observed score variance could then be thought of as true score variance plus error variance. The reliability of a test, rating scale, assessment or any other more or less standardized procedure within a given (sub)population of persons (or other objects of measurements, e.g. classrooms) is defined as the ratio of true score variance to observed score variance or as the squared correlation between observed scores and true scores (Lord & Novick, 1968: 61):

Its minimum value is zero, its maximum value one. As will be demonstrated, the definition is not very useful until we have defined precisely what we mean by ‘error’.
After the 1960s, Item Response Theory, IRT for short, became an influential approach in test theory. With IRT person parameters on a latent scale replace true scores. At first sight, there seems to be no place for reliability within the context of IRT. It can be demonstrated, however, that reliability is an important concept in the newer test theoretical approach also.
Reliability and Sources of Variation
When the length of a person is measured repeatedly, we notice small differences in the reading of the length: there is error in the measurements. The same is the case in measuring a person's characteristics in psychological testing. When an intelligence test would be administered to a person repeatedly, we would expect scores to vary: again there is measurement error. Unfortunately, the experiment of repeatedly testing a person with the same measurement instrument is seldom done; in practice we should expect memory effects. Instead, we could administer two tests meant to measure the same construct. Then a score difference might not only be due to chance fluctuations in item responses, but also to differences in content. Many more sources of variation can be thought of; for example, systematic fluctuation of responses over time. Sources of variance due to person characteristics can be classified as lasting or temporary, and lasting or specific. Further, there are factors affecting test administration and there is a category for variance not accounted for otherwise. Most of the sources of variation in responses might be regarded as a source of error variation, but the same sources might be regarded as sources of true variation, depending on the purpose of the test administrator. Let us give an example, mentioned by Stanley (1971: 366), who discusses the subject of sources of variation extensively. A person may be fatigued on the day of testing and this influences test performance. When our interest is to predict performances over some period, reliability would be consistency over time. When the intercorrelations among tests administered at the same session are studied, consistency at that session is relevant. So, the definition of error depends on the purpose of the investigator, and this should determine the choice of reliability coefficient(s).
...
- 1. Theory and Methodology
- Ambulatory Assessment
- Assessment Process
- Assessor's Bias
- Automated Test Assembly Systems
- Classical and Modern Item Analysis
- Classical Test Theory
- Classification (General, including Diagnosis)
- Criterion-Referenced Testing: Methods and Procedures
- Cross-Cultural Assessment
- Decision (including Decision Theory)
- Diagnosis of Mental and Behavioural Disorders
- Diagnostic Testing in Educational Settings
- Dynamic Assessment (Learning Potential Testing, Testing the Limits)
- Ethics
- Evaluability Assessment
- Evaluation: Programme Evaluation (General)
- Explanation
- Factor Analysis: Confirmatory
- Factor Analysis: Exploratory
- Formats for Assessment
- Generalizability Theory
- History of Psychological Assessment
- Intelligence Assessment through Cohort and Time
- Item Banking
- Item Bias
- Item Response Theory: Models and Features
- Latent Class Analysis
- Multidimensional Item Response Theory
- Multidimensional Scaling Methods
- Multimodal Assessment (including Triangulation)
- Multitrait-Multimethod Matrices
- Needs Assessment
- Norm-Referenced Testing: Methods and Procedures
- Objectivity
- Outcome Assessment/Treatment Assessment
- Person/Situation (Environment) Assessment
- Personality Assessment through Longitudinal Designs
- Prediction (General)
- Prediction: Clinical vs. Statistical
- Qualitative Methods
- Reliability
- Report (General)
- Reporting Test Results in Education
- Self-Presentation Measurement
- Self-Report Distortions (including Faking, Lying, Malingering, Social Desirability)
- Test Adaptation/Translation Methods
- Test User Competence/Responsible Test Use
- Theoretical Perspective: Cognitive
- Theoretical Perspective: Cognitive-Behavioural
- Theoretical Perspective: Constructivism
- Theoretical Perspective: Psychoanalytic
- Theoretical Perspective: Psychological Behaviourism
- Theoretical Perspective: Psychometrics
- Theoretical Perspective: Systemic
- Trait-State Models
- Utility
- Validity (General)
- Validity: Construct
- Validity: Content
- Validity: Criterion-Related
- 2. Methods, Tests and Equipment
- Adaptive and Tailored Testing
- Analogue Methods
- Autobiography
- Behavioural Assessment Techniques
- Brain Activity Measurement
- Case Formulation
- Coaching Candidates to Score Higher on Tests
- Computer-Based Testing
- Equipment for Assessing Basic Processes
- Field Survey: Protocols Development
- Goal Attainment Scaling (GAS)
- Idiographic Methods
- Interview (General)
- Interview in Behavioural and Health Settings
- Interview in Child and Family Settings
- Interview in Work and Organizational Settings
- Neuropsychological Test Batteries
- Observational Methods (General)
- Observational Techniques in Clinical Settings
- Observational Techniques in Work and Organizational Settings
- Projective Techniques
- Psychoeducational Test Batteries
- Psychophysiological Equipment and Measurements
- Self-Observation (Self-Monitoring)
- Self-Report Questionnaires
- Self-Reports (General)
- Self-Reports in Behavioural Clinical Settings
- Self-Reports in Work and Organizational Settings
- Socio-Demographic Conditions
- Sociometric Methods
- Standard for Educational and Psychological Testing
- Subjective Methods
- Test Accommodations for Disabilities
- Test Anxiety
- Test Designs: Developments
- Test Directions and Scoring
- Testing through the Internet
- Unobtrusive Measures
- 3. Personality
- Anxiety Assessment
- Attachment
- Attitudes
- Attribution Styles
- Big Five Model Assessment
- Burnout Assessment
- Cognitive Styles
- Coping Styles
- Emotions
- Empowerment
- Interest
- Leadership Personality
- Locus of Control
- Motivation
- Optimism
- Person/Situation (Environment) Assessment
- Personal Constructs
- Personality Assessment (General)
- Personality Assessment through Longitudinal Designs
- Prosocial Behaviour
- Self-Control
- Self-Efficacy
- Self-Presentation Measurement
- Self, The (General)
- Sensation Seeking
- Social Competence (including Social Skills, Assertion)
- Temperament
- Time Orientation
- Trait-State Models
- Values
- Weil-Being (including Life Satisfaction)
- 4. Intelligence
- Attention
- Cognitive Ability: g Factor
- Cognitive Ability: Multiple Cognitive Abilities
- Cognitive Decline/Impairment
- Cognitive Plasticity
- Cognitive Processes: Current Status
- Cognitive Processes: Historical Perspective
- Cognitive/Mental Abilities in Work and Organizational Settings
- Creativity
- Dynamic Assessment (Learning Potential Testing, Testing the Limits)
- Emotional Intelligence
- Equipment for Assessing Basic Processes
- Fluid and Crystallized Intelligence
- Intelligence Assessment (General)
- Intelligence Assessment through Cohort and Time
- Language (General)
- Learning Disabilities
- Memory (General)
- Mental Retardation
- Practical Intelligence: Conceptual Aspects
- Practical Intelligence: Its Measurement
- Problem Solving
- Triarchic Intelligence Components
- Wisdom
- 5. Clinical and Health
- Anger, Hostility and Aggression Assessment
- Antisocial Disorders Assessment
- Anxiety Assessment
- Anxiety Disorders Assessment
- Applied Behavioural Analysis
- Applied Fields: Clinical
- Applied Fields: Gerontology
- Applied Fields: Health
- Caregiver Burden
- Child and Adolescent Assessment in Clinical Settings
- Clinical Judgement
- Coping Styles
- Counselling, Assessment in
- Couple Assessment in Clinical Settings
- Dangerous/Violence Potential Behaviour
- Dementia
- Diagnosis of Mental and Behavioural Disorders
- Dynamic Assessment (Learning Potential Testing, Testing the Limits)
- Eating Disorders
- Health
- Identity Disorders
- Interview in Behavioural and Health Settings
- Irrational Beliefs
- Learning Disabilities
- Mental Retardation
- Mood Disorders
- Observational Techniques in Clinical Settings
- Outcome Assessment/Treatment Assessment
- Palliative Care
- Prediction: Clinical vs. Statistical
- Psychoneuroimmunology
- Quality of Life
- Self-Observation (Self-Monitoring)
- Self-Reports in Behavioural Clinical Settings
- Social Competence (including Social Skills, Assertion)
- Stress
- Substance Abuse
- Test Anxiety
- Thinking Disorders Assessment
- Type A: A Proposed Psychosocial Risk Factor for Cardiovascular Diseases
- Type C: A Proposed Psychosocial Risk Factor for Cancer
- 6. Educational and Child Assessment
- Achievement Testing
- Applied Fields: Education
- Child Custody
- Children with Disabilities
- Coaching Candidates to Score Higher on Tests
- Cognitive Psychology and Assessment Practices
- Communicative Language Abilities
- Development (General)
- Development: Intelligence/Cognitive
- Development: Language
- Development: Psychomotor
- Development: Socio-Emotional
- Diagnostic Testing in Educational Settings
- Dynamic Assessment (Learning Potential Testing, Testing the Limits)
- Evaluation in Higher Education
- Giftedness
- Instructional Strategies
- Interview in Child and Family Settings
- Item Banking
- Learning Strategies
- Performance
- Performance Standards: Constructed Response Item Formats
- Performance Standards: Selected Response Item Formats
- Planning
- Planning Classroom Tests
- Pre-School Children
- Psychoeducational Test Batteries
- Reporting Test Results in Education
- Standard for Educational and Psychological Testing
- Test Accommodations for Disabilities
- Test Directions and Scoring
- Testing in the Second Language in Minorities
- 7. Work and Organizations
- Achievement Motivation
- Applied Fields: Forensic
- Applied Fields: Organizations
- Applied Fields: Work and Industry
- Career and Personnel Development
- Centres (Assessment Centres)
- Cognitive/Mental Abilities in Work and Organizational Settings
- Empowerment
- Interview in Work and Organizational Settings
- Job Characteristics
- Job Stress
- Leadership in Organizational Settings
- Leadership Personality
- Motor Skills in Work Settings
- Observational Techniques in Work and Organizational Settings
- Organizational Culture
- Performance
- Personnel Selection, Assessment in
- Physical Abilities in Work Settings
- Risk and Prevention in Work and Organizational Settings
- Self-Reports in Work and Organizational Settings
- Total Quality Management
- 8. Neurophysiopsychological Assessment
- Applied Fields: Neuropsychology
- Applied Fields: Psychophysiology
- Brain Activity Measurement
- Dementia
- Equipment for Assessing Basic Processes
- Executive Functions Disorders
- Memory Disorders
- Neuropsychological Test Batteries
- Outcome Evaluation in Neuropsychological Rehabilitation
- Psychoneuroimmunology
- Psychophysiological Equipment and Measurements
- Visuo-Perceptual Impairments
- Voluntary Movement
- 9. Environmental Assessment
- Behavioural Settings and Behaviour Mapping
- Cognitive Maps
- Couple Assessment in Clinical Settings
- Environmental Attitudes and Values
- Family
- Landscapes and Natural Environments
- Life Events
- Organizational Structure, Assessment of
- Perceived Environmental Quality
- Person/Situation (Environment) Assessment
- Post-Occupancy Evaluation for the Built Environment
- Residential and Treatment Facilities
- Social Climate
- Social Networks
- Social Resources
- Stressors: Physical
- Stressors: Social
- Loading...
Get a 30 day FREE TRIAL
-
Watch videos from a variety of sources bringing classroom topics to life
-
Read modern, diverse business cases
-
Explore hundreds of books and reference titles
Sage Recommends
We found other relevant content for you on other Sage platforms.
Have you created a personal profile? Login or create a profile so that you can save clips, playlists and searches