Secondary Data in the ESS IA: Your Complete Student Guide

Hands analyzing environmental data sheets on desk

Secondary Data in the ESS IA: Your Complete Student Guide

Secondary data is fully allowed in the IB ESS Internal Assessment, and many of the strongest IAs are built entirely on it. The one rule that matters most: you must justify your dataset choice and critically evaluate the data using IB assessment language, not just download a spreadsheet and paste in a graph.

Before you commit to a secondary dataset, confirm three things quickly:

  • Relevance: Does the dataset directly address your research question (RQ)?
  • Provenance: Who collected the data, when, and how?
  • Citation and license: Can you legally reuse it, and do you have a citable DOI or URL?

The IB does not penalize you for using secondary data. What examiners do penalize is uncritical use: presenting data without questioning its limitations, without explaining why it suits your RQ, and without linking evidence to your conclusions. Get those three habits right, and secondary data can take you to a very high mark.

Pro Tip: Write one sentence in your Method section that explicitly states why secondary data is more appropriate for your RQ than primary collection would be. Examiners reward that justification under Criterion A.


Key Takeaways

Secondary data is a fully valid choice for the ESS IA when you justify the dataset, document your processing steps, and critically evaluate limitations in the language of IB criteria A–F.

Point Details
Secondary data is allowed IB ESS IA accepts secondary data; the key is justifying its use and evaluating it critically.
Justify under Criterion A State explicitly why secondary data is more appropriate than primary collection for your RQ.
Document every processing step Describe how you cleaned, subsetted, and derived variables; include a processing note in the Appendix.
Evaluate limitations specifically Name three precise limitations and connect each one to how it affects your conclusions.
Cite datasets with DOI or URL Include author, year, dataset title, publisher, and access date for every secondary source.

Table of Contents

When is secondary data acceptable for the ESS IA?

The ESS IA is a 1,500–2,250 word investigation assessed on six criteria: A (Focus and Research Question), B (Methodology), C (Results and Analysis), D (Discussion and Evaluation), E (Conclusion), and F (Presentation). Each criterion rewards evidence of genuine scientific thinking, not the method of data collection.

What counts as evidence in the IA:

  • Quantitative datasets (time series, monitoring records, survey data)
  • Qualitative secondary sources used to contextualize findings
  • Derived variables you calculate from raw secondary data (e.g., annual means, percentage change)
  • Figures and tables you construct from downloaded data

Explicit requirements when using secondary data:

  • State the original source clearly in your Method section
  • Explain the sampling design and collection method of the original dataset
  • Justify why this dataset is appropriate for your specific RQ
  • Acknowledge limitations of reusing data collected for a different purpose
  • Triangulate where possible: cross-check findings against a second source or a small primary dataset

Before you finalize your approach, ask your supervisor:

  1. Does my RQ require measurements I genuinely cannot take myself, or am I choosing secondary data for convenience?
  2. Is the dataset recent enough and spatially relevant to my RQ?
  3. Have I identified at least two limitations of this data that I can discuss critically?

Getting supervisor sign-off on these questions early saves you from a methodology rewrite later. For a broader look at how to structure your IA, Esstutor’s full guide walks through each section in detail.


Which research questions fit secondary data, and which ones don’t?

Secondary data works best when your RQ requires a scale, time span, or geographic reach that you cannot realistically achieve through personal data collection. It tends to be a poor fit when your RQ depends on site-specific precision or experimental manipulation.

Good fits for secondary data:

  • Historic trends (e.g., 30-year deforestation rates, decadal temperature anomalies)
  • Large-scale spatial patterns (national emissions inventories, regional biodiversity indices)
  • Long-term monitoring datasets where continuity matters
  • Sites that are physically inaccessible (remote ecosystems, international locations)
  • Socioeconomic indicators paired with environmental variables

Poor fits:

  • RQs requiring direct field measurements at a specific local site
  • Experimental designs where you need to control variables yourself
  • Questions about very recent or highly localized events with no existing dataset

Paired RQ examples:

Theme Poor fit (primary needed) Good fit (secondary data)
Urban heat islands “How does surface temperature vary across three city blocks in my town on one afternoon?” “How have urban heat island intensities changed in US cities over recent decades?”
Water quality “What is the current nitrate concentration in the stream behind my school?” “How have nitrate levels in US agricultural watersheds changed since the Clean Water Act?”
Biodiversity trends “How many bird species are present in my local park this spring?” “How has vertebrate species richness changed in North American protected areas since 2000?”

Environmental monitoring equipment in forest at dawn

The left-column RQs need your own measurements. The right-column RQs are genuinely unanswerable without long-term datasets, which is exactly the justification examiners want to see. For more guidance on choosing between data types, Esstutor’s data role guide is a good next read.

Pro Tip: Phrase your justification as a constraint, not a preference: “Primary data collection over the required 30-year period is not feasible within the scope of this investigation.” That framing directly satisfies Criterion A.


How to choose reliable secondary data: a practical checklist you can use

Choosing a dataset is a research decision, not a Google search. Examiners reading Criterion B want to see that you evaluated your source before using it. Run every candidate dataset through this checklist.

Data-quality checklist:

  • Authority: Is the data published by a government agency, research institution, or peer-reviewed journal? Avoid blogs, news aggregators, or unattributed spreadsheets.
  • Recency: Is the dataset current enough for your RQ? A 2005 emissions inventory is not suitable for a question about post-2015 trends.
  • Method transparency: Does the source describe how data were collected, by whom, and with what instruments or protocols?
  • Spatial resolution: Does the dataset cover the geographic scale your RQ requires (local, national, global)?
  • Temporal resolution: Are data points annual, monthly, or daily? Does that match what your analysis needs?
  • Metadata availability: Can you download a data dictionary or codebook that defines each variable and its units?
  • Sampling design: Was the original sampling random, stratified, or opportunistic? Opportunistic sampling introduces bias you must discuss.
  • Units: Are units consistent across years and countries? Mixed units (e.g., metric vs. imperial, CO₂ vs. CO₂-equivalent) require conversion and documentation.

Red flags to watch for:

  • No author or institutional affiliation listed
  • No date of collection or last update
  • No description of measurement methodology
  • Data that cannot be downloaded in a raw format (only pre-made charts)
  • Terms of use that prohibit academic reuse

Applying the checklist: a quick example. Suppose your RQ asks about global CO₂ emissions trends since 1990. You find a dataset on Our World in Data, which curates authoritative global datasets on climate, emissions, and land use and provides downloadable CSVs with full documentation. Running the checklist: the source is an established research publication (authority ✓), data runs to the most recent available year (recency ✓), methodology notes link to the original Global Carbon Project (method transparency ✓), data are available at country and global level (spatial resolution ✓), annual values are provided (temporal resolution ✓), a codebook defines each column (metadata ✓). That is a dataset worth using. For evaluating the advantages and limitations of secondary data more broadly, ATLAS.ti’s overview of primary versus secondary data is a useful reference.

Pro Tip: Paste your checklist responses into your Appendix. It takes five minutes and directly demonstrates the critical thinking Criterion B rewards.


Where should you search first for ESS secondary data?

The sources below are reliable, free, and well-documented. Each entry notes what type of ESS RQ it suits best. For a curated list with additional database links, Esstutor’s ESS data sources page is a practical starting point.

Global and cross-national sources:

  • Our World in Data: Climate change, CO₂ emissions, deforestation, energy use, biodiversity loss, and population. Downloadable CSVs with clear documentation. Best for long-term global trend RQs.
  • World Bank Open Data: Country-level time series on development, energy access, agriculture, and environmental indicators. Excellent for pairing socioeconomic and environmental variables across decades.
  • data.gov: The US federal open data portal. Covers air quality, water, land use, energy, and more. Best for US-focused RQs needing government-verified data.
  • Epa: US air quality index data, greenhouse gas emissions inventories, Superfund site records, and water quality monitoring. Ideal for US-based environmental policy RQs.
  • Usgs: Streamflow, groundwater levels, earthquake records, land cover change, and biodiversity surveys. Strong for water and land-system RQs.
  • Noaa: Climate normals, sea surface temperatures, hurricane tracks, sea level rise, and ocean chemistry. The go-to source for climate and ocean-related RQs.
  • LabWrite (NCSU): Not a dataset portal, but an essential methodological guide. It explains how primary measurements become secondary data and how to document data-processing steps in a methods section.

For US-focused IAs specifically: Start with EPA for air and water quality, USGS for hydrological and land-cover data, and NOAA for climate variables. Each portal lets you filter by state, watershed, or monitoring station, so you can make a national dataset locally relevant to your RQ.

Two sources to try first for most ESS topics: Our World in Data for global environmental trends and EPA for US-specific monitoring data. Both provide downloadable files, clear metadata, and are widely accepted in student research.


How to process and analyze secondary data for the IA

Downloading a dataset is the start, not the finish. What you do with the data determines your Criterion C score. Follow this workflow to produce clean, defensible results.

  1. Download and inspect the raw file. Open the CSV or spreadsheet and read every column header. Check for blank cells, inconsistent formatting, and mixed units before doing anything else.
  2. Read the metadata. The codebook or data dictionary tells you what each variable measures, the units used, and any known data gaps. Note any years or regions flagged as incomplete.
  3. Subset to your scope. Filter rows to the time range and geographic area your RQ specifies. Do not analyze data outside your stated scope; examiners notice when graphs include years your RQ does not address.
  4. Standardize units. Convert all values to a single consistent unit (e.g., all emissions in metric tons of CO₂-equivalent). Document every conversion in your Method section.
  5. Handle missing values explicitly. Do not delete missing rows silently. State in your Method how you treated gaps: excluded, interpolated, or replaced with the period mean. Each choice is a limitation to discuss later.
  6. Aggregate or normalize as needed. If your RQ asks about trends, compute annual means from monthly values. For example, to get the annual mean temperature from 12 monthly values, sum them and divide by 12. Show this calculation or describe it clearly so a reader could reproduce it.
  7. Construct your figures. Build graphs and tables from your processed data, not from screenshots of someone else’s chart. Your figure must be your own construction, even when the underlying numbers come from a secondary source.
  8. Write figure captions that explain what you did. A strong caption reads: “Figure 1: Annual mean CO₂ concentration (ppm) at Mauna Loa Observatory, 1990–2022. Data sourced from NOAA Global Monitoring Laboratory; monthly values averaged by the author to produce annual means.”

LabWrite emphasizes documenting every processing step so another researcher could replicate your analysis. A short processing note or script snippet in your Appendix satisfies that standard and signals methodological rigor to examiners. For students who want to practice the underlying calculations, MathVault offers clear numerical tutorials on data transformations and statistical operations.

Pro Tip: Never paste a screenshot of a graph from a website into your IA. Reconstruct the graph yourself using the raw data. That act alone demonstrates the data-handling competence Criterion C rewards.


Writing the IA with secondary data: how to align to IB criteria A–F

Every section of your IA needs to speak directly to the assessment criteria. Here is how secondary data maps to each one.

Criterion What examiners expect when secondary data is used
A: Focus and RQ A clear, focused RQ; explicit justification of why secondary data is appropriate for this specific question
B: Methodology Full description of the dataset’s origin, collection method, and sampling design; your processing steps documented
C: Results and Analysis Your own figures constructed from raw data; statistical or descriptive analysis linked to the RQ
D: Discussion and Evaluation Critical evaluation of data limitations; discussion of how limitations affect conclusions; triangulation where used
E: Conclusion A conclusion that directly answers the RQ and is supported only by evidence presented in C
F: Presentation Correct citation of all datasets; figures labeled and captioned; word count within limits

Examiner dos and don’ts:

  • Do name the dataset and its publisher in your Method section on the first mention.
  • Do explain the original purpose of the dataset and how that compares to your RQ.
  • Do use hedged language in your Analysis: “The data suggest…” rather than “The data prove…”
  • Don’t present a graph downloaded directly from a website as your own figure.
  • Don’t list limitations as a separate afterthought paragraph. Weave them into your Discussion.
  • Don’t ignore the gap between when data were collected and when your IA was written.

Common marking pitfalls with secondary-data IAs:

  • Criterion B scored low because the student described the dataset but not their own processing steps
  • Criterion D scored low because limitations were listed without explaining how each one affects the conclusions
  • Criterion F scored low because datasets were cited as website names only, with no author, date, or URL

Mini annotated example:

Method excerpt: "Annual mean surface temperature anomalies (°C, relative to the 1951–1980 baseline) were obtained from the NASA GISS Surface Temperature Analysis (GISTEMP v4) for the contiguous United States, 1980–2022. Monthly station data were averaged by the author to produce annual means for each decade.

Results caption: “Figure 2: Decadal mean surface temperature anomaly (°C) for the contiguous United States, 1980–2022. Data: NASA GISTEMP v4; processing by author.”

Analysis claim: “The data suggest a consistent warming trend across all four decades examined, with the 2010–2022 mean anomaly (+1.1°C) exceeding the 1980–1989 mean (+0.3°C) by 0.8°C. This pattern is consistent with broader Northern Hemisphere warming documented in the literature, though regional variability within the dataset limits generalization to specific states.”

For a deeper look at how each criterion is scored, Esstutor’s IA criteria guide breaks down the mark bands with practical examples.

Pro Tip: Read your Discussion paragraph aloud and ask: “Does every limitation I mention connect to a specific conclusion I drew?” If any limitation floats without a connection, cut it or link it. Examiners reward precision, not length.


Permissions, licenses, and citing datasets correctly in the ESS IA

Using someone else’s data without checking the license is an academic integrity risk, even in a student IA. Most government and research datasets are open, but the terms vary.

How to check a dataset’s license:

  • Look for a “Terms of Use,” “Data License,” or “Open Data” statement on the download page.
  • Many datasets use Creative Commons licenses, which specify whether you can reuse, adapt, and share the data and what attribution is required. A CC BY license requires attribution; CC BY-SA requires attribution and that any derivative work uses the same license.
  • US federal government data (EPA, USGS, NOAA, data.gov) is generally in the public domain under US law, but always confirm on the specific dataset’s page.
  • If no license is stated, contact the data owner before using the data.

Quick permissions email template:

Citation formats:

DOI/URL format (preferred when a DOI exists):

Author(s) or Organization. (Year). Dataset title [Data set]. Publisher. https://doi.org/xxxxx

Standard author-title-publisher format:

US Environmental Protection Agency. (2023). Air Quality System Data Mart [Data set]. EPA. Retrieved March 15, 2025, from https://www.epa.gov/aqs

Caption attribution template:
Pro Tip: Record the exact URL and access date the day you download the data. Dataset URLs sometimes change, and examiners need enough information to locate the original source.


How to discuss validity, reliability, and bias when your IA uses secondary data

Critical evaluation of your data is where many students lose marks on Criterion D. The goal is not to apologize for using secondary data. It is to show you understand exactly how its limitations shape your conclusions.

Language templates for common limitations:

  • Sampling bias: “The dataset relies on monitoring stations concentrated in urban areas, which may overrepresent pollution levels relative to rural regions and limit the generalizability of findings to the broader study area.”
  • Temporal mismatch: “The most recent data available are from 2021, meaning trends in the final two years of the study period cannot be assessed and may affect the conclusion regarding recent acceleration.”
  • Measurement units: “Emissions figures prior to 2005 are reported in CO₂ equivalents using the IPCC Second Assessment Report global warming potentials, while post-2005 values use Fourth Assessment Report values, introducing a small systematic inconsistency.”
  • Proxy variables: “NDVI (Normalized Difference Vegetation Index) is used here as a proxy for vegetation cover, but it cannot distinguish between species types, which limits conclusions about biodiversity specifically.”

Bias checklist:

  • Was the data collected for a purpose different from your RQ? If so, how might that affect what was measured?
  • Are there geographic regions or time periods underrepresented in the dataset?
  • Could the measurement instrument or protocol have changed over the study period?
  • Does the dataset aggregate data in ways that mask variation relevant to your RQ?
  • Are there political or funding incentives that could have influenced what was reported?

On triangulation: Combining a small primary dataset with a larger secondary one is one of the most effective ways to strengthen an ESS IA. For example, you might take three local air quality readings with a portable sensor and compare them to the nearest EPA monitoring station’s annual averages. That comparison does not require a large primary dataset; it just requires you to discuss what the agreement or disagreement between the two sources tells you. Even a brief triangulation note in your Discussion signals methodological awareness to examiners.

Pro Tip: Aim for three specific, well-explained limitations rather than six vague ones. “The data may not be fully accurate” earns nothing. “The dataset excludes informal settlements, which likely have higher exposure rates and would shift the mean upward” earns marks.


A compact annotated mini-IA plan using secondary data

Here is a one-page plan you can copy and adapt for your own investigation.

Research question: How have annual mean PM2.5 concentrations changed in the 10 most populous US cities between 2000 and 2022?

Justification for secondary data: Long-term air quality monitoring requires continuous instrumentation over decades, which is not feasible within the scope of a student IA. The EPA Air Quality System (AQS) provides validated, station-level PM2.5 data for this period.

Dataset: EPA Air Quality System Data Mart (epa.gov/aqs). Annual summary files for PM2.5 (FRM/FEM), filtered to the 10 cities by CBSA code. Data are in the public domain.

Methods summary:

  1. Download annual summary CSV files for PM2.5, 2000–2022, from EPA AQS.
  2. Filter to monitoring stations within the 10 target cities.
  3. Calculate annual mean PM2.5 for each city by averaging all valid station readings per year.
  4. Construct a line graph showing annual mean PM2.5 (µg/m³) for each city over time.
  5. Calculate the percentage change in mean PM2.5 from 2000 to 2022 for each city.

Expected analysis: Trend analysis comparing decadal means; identification of cities with the largest and smallest reductions; discussion of whether trends align with Clean Air Act regulatory milestones.

Conclusions to test: Whether PM2.5 concentrations have declined consistently across all cities, or whether improvement is uneven and related to regional policy differences.

Appendix checklist for submission:

  • [ ] Metadata printout from EPA AQS (dataset description, variable definitions, units)
  • [ ] DOI or permanent URL with access date
  • [ ] Confirmation of public domain status (EPA data use statement)
  • [ ] Processing steps described in Method or as a numbered note in the Appendix
  • [ ] Raw data file included or referenced with retrieval instructions
  • [ ] All figures labeled, captioned, and constructed by the author from raw data
  • [ ] Limitations section addresses at least three specific constraints of the AQS dataset
  • [ ] Supervisor sign-off on dataset choice and methodology before final submission

For additional RQ ideas mapped to specific databases, Esstutor’s ESS IA topics and secondary databases page lists options organized by ESS topic area.


What most students get wrong about secondary data in the ESS IA

Most students treat secondary data as the easy option. They assume that because the numbers already exist, the hard work is done. That assumption is exactly what costs marks.

The real challenge with secondary data is not finding it. It is demonstrating that you understand it well enough to question it. An examiner reading a Criterion D response wants to see a student who knows why the EPA’s monitoring network underrepresents rural communities, or why a global land-use dataset might classify degraded forest as intact canopy depending on the resolution used. That kind of specific, source-aware critique is what separates a 5 from a 7.

There is also a tendency to over-rely on one dataset and treat it as ground truth. The strongest IAs I see use a primary source, even a small one, to anchor the secondary data. Three local measurements compared to a national average tells an examiner you understand the difference between what a dataset claims and what is actually happening at your study site. That comparison, even when the results are inconclusive, is methodologically honest in a way that a single-source IA rarely is.

One more thing: the word limit is 1,500–2,250 words. Students using secondary data often spend too many words describing the dataset and too few analyzing what it shows. Flip that ratio. A two-sentence dataset description followed by three paragraphs of analysis is far more likely to score well than the reverse.

If you want structured support working through your dataset choice, methodology, and evaluation language, Esstutor offers one-on-one ESS IA tutoring with an IB examiner who has guided students through exactly this process for over 13 years.

Sources

Use this list as your first stop when searching for data. Include dataset metadata and license information in your Appendix for every source you use.

Reminder: Every dataset you use should have its metadata (variable definitions, units, collection method, date range) and license information printed or saved and included in your IA Appendix. Examiners cannot verify what they cannot find.


No Comments

Post A Comment