Sources of Crime and Criminal Justice Data
Most criminological research reuses data already collected by police, courts, correctional agencies or earlier surveys rather than gathering it fresh. Knowing where such data comes from, and where it breaks down, is a precondition for using it responsibly.
Criminological research relies overwhelmingly on data that someone else collected first. Sources of crime and criminal justice data are the channels through which that information reaches a researcher: administrative records kept by police, courts and prisons for operational reasons, national statistical series compiled from those records, purpose-built cohort and survey datasets, and increasingly, digital trace data left behind by online activity.
Each source was built for a purpose other than the research question a criminologist eventually asks of it. A police record exists to support an investigation and a prosecution, not to measure the true rate of an offence in a population.
A national statistical series exists to inform policy and resourcing at the aggregate level, not to trace an individual life course. Recognising the original purpose of a dataset is the first step in judging what it can and cannot answer.
This topic maps the main families of data researchers actually use: administrative and official statistics, longitudinal cohort datasets, secondary survey data, linked administrative records and emerging digital trace sources. It also sets out the quality problems, missingness and cross-jurisdiction incomparability, that follow every one of them into an analysis.
By the end of this topic you should be able to:
- Distinguish primary data collection from secondary use of an existing dataset, with examples of each in criminology.
- Explain what administrative records from police, courts and corrections can and cannot support as research evidence.
- Describe how a birth-cohort or offender-cohort study differs in design from a cross-sectional survey.
- Explain what record linkage does, and name at least one ethical or technical constraint it faces.
- Identify at least two threats to data quality that affect comparison of crime data across jurisdictions.
- Administrative data
- Records generated by an agency, such as police, courts or prisons, as a by-product of its operational work, later reused for research or statistics.
- Secondary data
- Data collected by someone else for a different purpose and reanalysed by a researcher asking a new question of it.
- Cohort study
- A longitudinal design that follows the same defined group of people, typically sharing a birth year or an event, over an extended period of time.
- Record linkage
- The process of matching records belonging to the same individual across two or more separate datasets, using shared identifiers or probabilistic matching.
- Open data
- Government or institutional data released in a structured, machine-readable format for public download and reuse, usually through a dedicated portal.
- Dark figure of crime
- The gap between the true volume of offending in a population and the volume that appears in official records.
Primary versus secondary data in criminological research
Primary data is collected directly by the researcher to answer a specific question: a new victimisation survey fielded in three cities, or a set of structured interviews with people recently released from prison. The researcher controls the sampling frame, the question wording and the timing, which means the data fits the research question closely, but collecting it is slow and expensive, and it usually covers a narrow slice of time and place.
Secondary data was collected by an agency or an earlier study for a purpose that predates the researcher's question. Most criminology relies on it, because building a new national dataset from scratch is rarely feasible for a single project. A researcher studying sentencing disparity, for instance, is far more likely to obtain court case files or a court administration extract than to interview every judge in a jurisdiction.
The tradeoff is fit versus reach. Secondary sources cover large populations and long time spans that no single research grant could fund directly, but the categories, definitions and missing fields were fixed by whoever built the dataset first, not by the later researcher's hypothesis. A police record that does not distinguish attempted from completed burglary, because operational practice never needed that distinction, cannot be made to distinguish them after the fact.
Most published criminological work sits on a spectrum rather than at either pole: a study might combine a secondary administrative extract with a small primary interview sample added to explain patterns the administrative data cannot interpret on its own. Recognising which parts of a study's evidence are primary and which are secondary is basic groundwork before assessing its conclusions.
Official administrative data: police, court and correctional records as research inputs
Police records are the most widely used administrative source in criminology. They typically include an incident report, an offence classification, and, where relevant, details of an arrest.
In the United States, most police agencies now report incident-level data to the FBI's National Incident-Based Reporting System (NIBRS), which replaced the older summary-based Uniform Crime Reporting (UCR) count as the national collection standard. In England and Wales, the Home Office compiles police-recorded crime data that the Office for National Statistics publishes alongside the independent Crime Survey for England and Wales.
Court records add a second layer: charge sheets, case dispositions, sentencing outcomes and appeal records. These let researchers study case attrition, the drop-off between an initial charge and a final conviction, and sentencing patterns across offence types, courts or demographic groups. Correctional records, covering admissions, time served, disciplinary incidents and release conditions, support research on incarceration trends and reoffending after release.
In India, the National Crime Records Bureau (NCRB) has compiled the annual Crime in India report since 1953, drawing on state police returns under offence categories that now sit under the Bharatiya Nyaya Sanhita, 2023, which came into force on 1 July 2024 and replaced the Indian Penal Code, 1860 as the general criminal code.
Researchers using NCRB series that span the transition need to track which offence definitions changed at the switchover, because a category-to-category comparison across the two codes is not automatic.
The recurring limitation across all three record types is selection. A case only enters a court dataset if police made an arrest and a prosecutor filed a charge; it only enters a correctional dataset if the case ended in a custodial disposition. Each stage filters out cases for reasons that are themselves criminologically interesting, discretion, plea bargaining, resource constraints, so administrative data describes the output of the justice system's decisions at least as much as it describes the underlying offending.
National statistical series and open crime-data portals
Above individual agency records sits a layer of aggregated national series designed for public reporting and cross-time comparison. The FBI's Crime Data Explorer publishes NIBRS-derived statistics for United States law enforcement agencies; the Office for National Statistics publishes the UK's Crime Survey for England and Wales alongside police-recorded crime; the Australian Bureau of Statistics publishes the annual Recorded Crime - Victims collection; and NCRB's Crime in India report performs the equivalent role for India. Each is built by aggregating agency-level administrative data upward, so it inherits every classification and reporting limitation already present at the source.
Open government data portals have made a growing share of this material directly downloadable rather than locked in periodic PDF reports. Portals such as data.gov (United States), data.gov.uk and India's data.gov.in publish structured datasets, sometimes at a finer geographic resolution than the flagship annual report, letting a researcher build custom local-area or multi-year extracts without submitting a formal data request.
Open portals still carry the same upstream problems as the source records: what is not recorded at the police station is not recorded in the portal either. What they add is transparency about methodology and, in the better-run examples, machine-readable metadata describing exactly how a count was constructed, which is often the only way a researcher can tell whether two years of a series are actually comparable.
National series are also where methodology changes are most visible, and most disruptive to a time series. A change in counting rules, an expanded set of reportable offences, or a shift from summary to incident-based reporting (as with the UCR to NIBRS transition in the United States) can produce a jump or dip in a published rate that reflects the change in method rather than a change in crime. Reading the accompanying technical notes before using a series across a methodology change is not optional.
Longitudinal and cohort datasets: birth-cohort and offender-cohort studies
A cohort study follows a defined group over time rather than sampling a fresh cross-section at each wave. In criminology this usually means either a birth cohort, everyone born in a given place and year, or an offender cohort, everyone who came into contact with the justice system in a given period.
Because the same individuals are tracked repeatedly, cohort data supports questions that a single-wave survey cannot answer directly: at what age does offending typically begin, how long does it persist, and what precedes desistance.
The Cambridge Study in Delinquent Development, begun by Donald West in 1961 and later directed by David Farrington, followed 411 London boys from age eight into adulthood, and remains one of the longest-running criminological cohort studies in the world.
In the United States, Marvin Wolfgang, Robert Figlio and Thorsten Sellin's 1972 study Delinquency in a Birth Cohort tracked boys born in Philadelphia in 1945 and introduced the influential finding that a small share of a cohort accounts for a disproportionate share of its recorded offending, a pattern later described as the chronic offender effect.
Cohort designs carry their own costs. Attrition, cohort members who move, decline further contact or cannot be traced, accumulates with every wave and can bias later findings toward whoever remained easiest to follow. Funding a study across decades is also rare; most cohort studies that have produced durable findings ran for far longer than a typical grant cycle, which is part of why so few exist and why the handful that do are cited so heavily.
Offender cohorts, built from justice-system contact rather than birth records, are cheaper to assemble because the sampling frame already exists in agency data, but they inherit the selection problem described in section 2: an offender cohort defined by arrest already excludes everyone whose offending never resulted in an arrest, which limits what it can say about the causes of offending onset in the wider population.
Survey data as a secondary source: reusing victimisation and self-report datasets
Victimisation surveys, such as the National Crime Victimization Survey in the United States or the Crime Survey for England and Wales, and self-report offending surveys ask a sample of the population directly about crimes they experienced or committed, independent of whether those incidents were ever reported to police. These surveys were designed to estimate the dark figure of crime, but once fielded, the resulting dataset becomes a reusable secondary resource for questions the original survey team never asked.
A researcher studying fear of crime, for example, might reanalyse several waves of an existing victimisation survey rather than fielding a new one, because the survey already contains relevant questions, a large sample and a known sampling design. This secondary reuse is efficient, but it locks the researcher into whatever variables, wording and sampling frame the original instrument used; a question that was not asked in the original survey cannot be recovered from the dataset later.
Repeated cross-sectional surveys, run on a regular cycle with a fresh sample each time, let researchers track national trends over time even though they do not follow the same individuals, which is the key structural difference from the cohort designs discussed in section 4. Comparing a repeated cross-sectional trend against a police-recorded trend for the same period is one of the standard ways researchers estimate how the dark figure of crime is moving, not just its current size.
Self-report surveys carry a distinct measurement concern: respondents may under-report socially sensitive or illegal behaviour, or occasionally over-report it, and the direction and size of that bias can differ by offence type and by who is asking the question. Researchers reusing an existing self-report dataset should treat its published validation work, where the survey organisation tested question wording against known outcomes, as part of the dataset's documentation, not as optional background reading.
Record linkage: matching individuals across agency datasets
Record linkage combines two or more separate administrative datasets by matching the records that belong to the same individual, using either a shared unique identifier or probabilistic matching on fields such as name, date of birth and address. The technique originates outside criminology, in Halbert Dunn's 1946 proposal for linking vital-statistics and health records, and has since become a standard tool wherever a research question needs a fuller picture than any single agency's records provide on their own.
In criminology, linkage is used to connect justice-system records with health, education, employment or welfare records, which lets researchers ask questions no single-agency dataset can answer: whether a childhood contact with child protective services predicts later justice-system contact, or whether access to employment support after release reduces reoffending. Because each linked dataset was built by a different agency, on a different schedule and with different data standards, linkage projects typically need substantial data-cleaning work before matching can even begin.
Linkage raises privacy and consent questions that a single-agency dataset does not, because a linked record can reveal far more about an individual than any one source alone, and the individual concerned may never have consented to that combination being created. Jurisdictions that support large-scale linkage for research, several university-run data linkage centres in Australia and the United Kingdom are established examples, generally operate under strict de-identification, secure-access and ethics-review requirements precisely because of this added disclosure risk.
Match quality is also a technical limit on its own. Probabilistic linkage produces both false matches, two different people wrongly treated as one, and missed matches, one person's records wrongly treated as belonging to two people, and either error can distort downstream findings. Reporting a linkage study's match rate and validation approach is treated as basic methodological disclosure, not an optional footnote.
Big data and digital trace data: promise and limits for criminology
Digital trace data refers to records generated automatically as a by-product of online or networked activity: social media posts, mobile phone location logs, online marketplace transactions, or server logs from a platform used to plan or commit an offence. Unlike administrative or survey data, no one designed a collection instrument in advance; the researcher captures activity that was already happening for a different reason, closer in spirit to observing a natural record than to fielding a survey.
The appeal is scale and immediacy. A dataset of millions of social media posts or transactions can be assembled far faster than a national survey, and it can capture behaviour, such as online fraud, cyberbullying or the sale of illicit goods on dark-web marketplaces, that conventional victimisation surveys were never designed to ask about. Cyber-forensics research in particular has drawn on platform and network logs that have no equivalent in traditional crime statistics.
The limits are just as significant. Digital trace datasets are rarely representative of any defined population, because access to the platform that generated them, and willingness to use it in a way that leaves a visible trace, are themselves unevenly distributed.
Platform terms of service, data-protection law and, in the case of law-enforcement-obtained data, evidentiary rules also restrict what a researcher can access and how it can be shared, which makes replication by an independent team harder than with a public statistical series.
Digital trace data is best treated as a complement to administrative and survey sources rather than a replacement for them. It can reveal patterns, in timing, network structure or the language used in fraudulent communication, that would never surface in an aggregate crime count, but establishing that a pattern in a platform log generalises beyond that platform's particular user base requires the same care that any non-random secondary sample demands.
Data quality, missingness and comparability across jurisdictions
Every source discussed above shares a common vulnerability: the quality of a secondary dataset is fixed by decisions the researcher had no part in. Missing fields, inconsistent classification over time, and gaps where an agency simply did not record a variable are the norm rather than the exception, and they rarely announce themselves in a summary table.
Cross-jurisdiction comparison compounds the problem. Offence definitions differ: what one country's statute treats as a single aggravated offence, another may record as two separate lesser offences. Counting rules differ too, some agencies count every offence within an incident, others count only the most serious offence per incident, which alone can produce very different totals from an identical set of underlying events.
India's 2023 to 2024 transition from the Indian Penal Code and Code of Criminal Procedure to the Bharatiya Nyaya Sanhita and Bharatiya Nagarik Suraksha Sanhita is a concrete example of a definitional break that a naive before-and-after comparison of NCRB figures would need to account for explicitly.
Reporting practices vary further with public trust in the police, awareness of how to report an offence, and cultural norms around disclosing certain crimes, particularly sexual and domestic offences. Two countries with an identical underlying rate of a sensitive offence can show very different recorded rates purely because victims in one context are more willing to report it, which is one reason victimisation surveys are treated as a necessary check on recorded-crime comparisons rather than an optional extra.
Good practice for handling these limits is not exotic: read the technical and methodological notes that accompany a series before comparing it across time or place, treat any comparison spanning a definitional or statutory change as provisional until the break is explicitly modelled, and report missingness rates rather than silently dropping incomplete cases. None of this eliminates the underlying data-quality problem, but it keeps a study's conclusions honest about what the data can actually support.
A dataset built by a court for case management, later reused by a researcher studying sentencing patterns, is an example of what kind of data?
Key Takeaways
- Most criminological research uses secondary data, collected by an agency or an earlier study for a different purpose than the one the later researcher is asking.
- Police, court and correctional records describe the justice system's decisions as much as they describe underlying offending, because each stage filters cases through discretion and attrition.
- National statistical series aggregate administrative data upward and inherit its classification limits, so methodology changes in a series can produce jumps that reflect counting rules, not crime.
- Cohort studies, such as the Cambridge Study in Delinquent Development and the Philadelphia birth-cohort research, track the same individuals over time and can study onset, persistence and desistance directly, at the cost of decades-long funding and attrition.
- Record linkage matches individuals across agency datasets to answer questions no single source can, but it raises privacy concerns and match-quality risks that single-agency data does not.
- Digital trace data offers scale and access to behaviour conventional sources miss, but it is rarely representative of a defined population.
- Comparing crime data across jurisdictions or across a statutory change, such as India's 2023 to 2024 shift to the Bharatiya Nyaya Sanhita, requires checking definitions and counting rules before treating the numbers as comparable.
What is the difference between primary and secondary crime data?
Why do police-recorded crime figures undercount actual crime?
What is a cohort study in criminology?
What is record linkage and why is it useful?
Can crime data be compared across different countries?
Test yourself on Criminology with free, timed mocks.
Practice Criminology questions