15 Aug How to Write a High-Scoring ESS IA Evaluation
A high-scoring ESS IA evaluation does one thing above all else: it links every strength, limitation, and improvement directly to your research question and your method. That connection is what separates a 6 from a 3 on Criterion F. Before you write a single sentence, check that you can answer this: “How does this point affect the validity or reliability of my specific investigation?” If you can’t, the point doesn’t belong.
Here is what every evaluation must include:
- Link to RQ and method — every point must name the method and explain how it affects your specific findings
- 2–4 strengths — explain how each one supported validity or reliability (not just “I used a control group”)
- 2–4 limitations — describe the likely direction and size of the effect, not just that an error existed
- Realistic improvements — explain how you would implement them and why they are feasible in a school setting
- 1–2 logical extensions — new research questions your data raises
- Criterion-weighted conclusion — name the decisive criterion (e.g., reliability) and give a final judgment
Sentence stems you can use right now:
- “The use of [method] strengthened the validity of the results because…”
- “A key limitation was [specific issue], which likely caused an underestimation of [variable] because…”
- “To improve reliability, [specific change] could be implemented by [how], which is feasible within a school context because…”
Quick example in use:
“The use of quadrat sampling strengthened the validity of the results because it reduced observer selection bias. A key limitation was the 1 m² quadrat size, which likely caused an underestimation of species diversity in dense vegetation because smaller organisms were missed. To improve reliability, a 4 m² quadrat could be used in future studies, which is feasible within a school context because no additional equipment is required.”
Pro Tip: Write your evaluation after re-reading your research question out loud. Every sentence you write should be answering: “What does this mean for my RQ?”
Key Takeaways
A high-scoring ESS IA evaluation links every limitation and improvement explicitly to the research question and method, using criterion-based paragraphs rather than unstructured lists.
| Point | Details |
|---|---|
| Link every point to the RQ | Each strength, limitation, and improvement must name the method and explain its effect on your specific findings. |
| Use criterion-based paragraphs | Structure each body paragraph around one criterion (validity, reliability, etc.) for higher Criterion F marks. |
| Make limitations directional | State whether each limitation caused an over- or underestimation, and explain why, to earn top-band marks. |
| Confirm improvement feasibility | Every proposed improvement needs a one-sentence note confirming it is achievable in a school setting. |
| Esstutor IA sessions | Marija reviews your evaluation paragraph by paragraph, identifying lost RQ links and mark-band gaps. |
Table of Contents
- What do IB examiners actually look for in the evaluation?
- How should you structure the evaluation paragraph by paragraph?
- Sentence stems and phrasing you can use in your evaluation
- What mistakes do students most often make in the evaluation?
- Annotated exemplar: what a strong evaluation paragraph looks like
- Your pre-submission checklist for the evaluation section
- A tutor’s perspective on the one edit that changes everything
- How Esstutor can help you polish your evaluation
- Sources
What do IB examiners actually look for in the evaluation?
The IB IA is marked across Criteria A–F, and the evaluation section maps directly to Criterion F. Vague statements earn the lowest mark band regardless of how many points you list.
Here is how the criteria connect:
| Criterion | Focus | How the evaluation affects it |
|---|---|---|
| A — Focus and research question | Clarity of the RQ | A weak evaluation that ignores the RQ signals the RQ was unclear from the start |
| D — Analysis and interpretation | Data interpretation | Limitations that explain anomalies in your data strengthen Criterion D too |
| E — Conclusion | Judgment tied to evidence | A strong conclusion in the evaluation reinforces the conclusion in Criterion E |
| F — Evaluation | Limitations, improvements, extensions | The primary criterion; marks depend on specificity and links to RQ/method |
Examiners use a best-fit model, meaning they look at the overall quality of your evaluation, not a checklist. A single precise, well-explained limitation outweighs three vague ones. The mark-band descriptors for the top band consistently reward criterion-based judgments and explicit links to the method and RQ, not generic observations.
How should you structure the evaluation paragraph by paragraph?
A criterion-led structure consistently outperforms unstructured pro/con lists in IB-style marking. Here is a paragraph plan that works:
-
Opening judgment (1–2 sentences): State a provisional verdict on whether your method was appropriate for your RQ. Use conditional language: “Overall, the method was largely appropriate for investigating X, though reliability was limited by Y.”
-
Body paragraph 1 — Strength: Name one strength, link it to the method, and explain its effect on validity or reliability. Avoid generic praise. “The use of a control site increased the internal validity of the comparison because it isolated the effect of [variable] from confounding land-use differences.”
-
Body paragraph 2 — Limitation with effect: Name one specific limitation, explain why it occurred given your method, and state the likely direction of the effect on your results. This is where most students lose marks. “The 48-hour sampling window may have underestimated nocturnal species activity, leading to a systematic undercount of [organism] in the treatment site.”
-
Body paragraph 3 — Improvement and feasibility: Propose one realistic change, explain how it would be implemented, and confirm it is achievable in a school setting. “Extending sampling to a 7-day period using automated data loggers, available through the school’s science department, would improve temporal reliability without requiring additional fieldwork days.”
-
Closing judgment (2–3 sentences): Name the decisive criterion, give a final verdict, and suggest one logical extension. Do not introduce new evidence here. Academic writing guides from the University of Göttingen confirm that conclusions should summarize without adding new claims, and the same principle applies to your IA.
Choosing your evaluation criteria: Pick 2–3 from this list based on what your method actually used: validity, reliability, sampling bias, instrumentation accuracy, spatial or temporal scale, ethical constraints, or uncertainty in measurements. Don’t choose criteria that don’t apply to your investigation.
Dos and don’ts:
- ✅ Name the specific instrument or technique that caused the limitation
- ✅ State whether the effect likely caused an over- or underestimation
- ✅ Confirm feasibility of improvements in a school context
- ❌ Write “human error” without specifying what, why, and in which direction
- ❌ List improvements that require university-level equipment or funding
- ❌ End with a new claim or piece of evidence in the conclusion
Pro Tip: After writing each limitation, ask yourself: “Would this limitation exist if I had used a different method?” If yes, name that method in your improvement. That connection is exactly what examiners reward.
Sentence stems and phrasing you can use in your evaluation
These stems are organized by function. Adapt them to your specific investigation. Where German is your working language, the German versions are provided first with English notes.
Opening judgment:
- “Die Methode war insgesamt geeignet, um die Forschungsfrage zu beantworten, obwohl die Zuverlässigkeit durch [X] eingeschränkt wurde.” (The method was overall appropriate for answering the RQ, though reliability was limited by [X].)
- “Overall, the investigation was valid for comparing [X] across [sites/conditions], with the primary limitation being [Y].”
Explaining a strength:
- “Die Verwendung von [Methode] erhöhte die Validität, da sie [Störvariable] kontrollierte.” (The use of [method] increased validity because it controlled for [confounding variable].)
- “The replicated sampling design strengthened reliability because it reduced the influence of random variation on the mean values.”
Stating a limitation and its effect:
- “Eine wesentliche Einschränkung war [X], was wahrscheinlich zu einer Überschätzung von [Variable] geführt hat, weil…” (A key limitation was [X], which likely caused an overestimation of [variable] because…)
- “The [instrument] had a measurement uncertainty of ±[value], which may have masked small differences between [groups], reducing the sensitivity of the comparison.”
Proposing an improvement:
- “Um die Zuverlässigkeit zu verbessern, könnte [konkrete Änderung] eingesetzt werden. Dies ist im schulischen Rahmen realisierbar, da [Begründung].” (To improve reliability, [specific change] could be used. This is feasible in a school context because [reason].)
- “Increasing the sample size from [n] to [n+] would reduce sampling error and is achievable within a standard school fieldwork day.”
Suggesting an extension:
- “Eine logische Erweiterung wäre die Untersuchung von [neuer Variable], um zu klären, ob [Hypothese].” (A logical extension would be to investigate [new variable] to determine whether [hypothesis].)
- “Future research could examine whether [finding] holds across different seasons, addressing the temporal limitation identified above.”
Criterion-weighted conclusion:
- “Der entscheidende Faktor für die Bewertung dieser Untersuchung ist die [Validität/Zuverlässigkeit], da [Begründung]. Insgesamt…” (The decisive criterion for evaluating this investigation is [validity/reliability] because [reason]. Overall…)
- “The most significant constraint on the conclusions was [criterion], because [explanation]. A future study addressing [improvement] would substantially strengthen the findings.”
A note on tone: Write in an objective, impersonal register. You may use “I” sparingly when describing a deliberate methodological choice (“I chose quadrat sampling because…”), but avoid first-person for limitations and improvements. Keep sentences concise. Examiners read hundreds of evaluations; clarity earns marks faster than elaborate phrasing. The University of Oldenburg’s essay rubric confirms that formal correctness and clear argument structure directly affect marks, and the same applies to IA writing.
What mistakes do students most often make in the evaluation?
These are the errors that cost the most marks, and each one has a direct fix:
-
Writing “human error” without specifics. Fix: Name the exact action (e.g., “inconsistent timing of water temperature readings”), explain why it happened, and state the likely direction of the bias. See the common IB ESS IA mistakes guide for more examples.
-
No link to the RQ or method. Fix: After every limitation sentence, add: “This affected the RQ because…” If you can’t complete that sentence, the point is too generic.
-
Pro/con lists with no criteria. Fix: Replace the list format with criterion-based paragraphs. Each paragraph should name one criterion (validity, reliability, etc.) and build an argument around it.
-
Unsupported claims about effect size. Fix: Use directional language (“likely caused an overestimation”) and bound the effect where possible (“by approximately [range]”). Don’t claim certainty you don’t have.
-
Unrealistic improvements. Fix: Every improvement must be achievable in a school setting. Avoid suggestions that require specialized lab equipment, large budgets, or multi-year timelines unless your school genuinely has access to them.
-
Overly informal or hedged tone. Fix: Avoid phrases like “I think,” “maybe,” or “it could possibly be.” Use “likely,” “probably,” “may have” for genuine uncertainty, but pair them with a reason.
-
Missing feasibility notes. Fix: Every improvement needs a one-sentence feasibility statement. Examiners want to know you understand the constraints of school-based research.
Annotated exemplar: what a strong evaluation paragraph looks like
Fictional context: Research question: “How does distance from a road (0 m, 25 m, 50 m) affect the species richness of ground-level vegetation in a temperate deciduous forest in Lower Austria?” Method: quadrat sampling (1 m² quadrats, n=10 per distance zone, single visit in April).
Exemplar paragraph:
“Overall, the quadrat sampling method was appropriate for comparing species richness across distance zones, though the reliability of the results was constrained by the single-visit design. [1] The standardized quadrat size and fixed transect positions increased the internal validity of the comparison by controlling for habitat heterogeneity between zones. [2] However, the April sampling window captured only spring ephemerals, likely causing an underestimation of total species richness at all three distances and potentially masking seasonal differences between zones. [3] This limitation directly affects the RQ because any difference in richness observed may reflect phenological variation rather than road proximity. [4] To improve temporal reliability, sampling could be repeated across three seasons (spring, summer, autumn) using the same quadrat positions; this is feasible within a school year and requires no additional equipment. [5] A logical extension would be to investigate whether heavy metal concentrations in soil (measurable with portable XRF analyzers available through university loan programs in Austria) correlate with the species richness gradient observed, which would help determine whether the road effect is chemical, physical, or both. [6] The decisive criterion limiting the conclusions is reliability, as the single-visit design prevents confirmation that the observed pattern is stable over time.” [7]
Inline annotations:
- Opening provisional judgment names the method and the decisive criterion
- Strength linked explicitly to validity and the specific design choice
- Limitation with direction of effect (underestimation) and mechanism
- Explicit link back to the RQ — this is what earns the mark
- Improvement with how, why, and feasibility confirmed for a school in Austria
- Logical extension with a new RQ and a realistic tool reference
- Criterion-weighted conclusion with no new evidence introduced
The IB ESS IA teacher guide notes that teachers may give one round of written feedback before final submission. Use that feedback specifically to strengthen the links between your limitations and your RQ — that is where the most marks are recoverable.

Your pre-submission checklist for the evaluation section
Run through this before you submit:
- Every limitation is linked to the RQ and method — not just named, but explained in terms of its effect on your specific findings.
- At least two criteria are named (e.g., validity and reliability) and each has its own paragraph or clear argument.
- Every improvement is realistic and testable in a school setting, with a one-sentence feasibility note.
- The conclusion names the decisive criterion and gives a final judgment without introducing new evidence.
- No sentence says “human error” without specifying what, why, and in which direction.
- Citations and referencing are correct — check the ESS IA referencing guide if you’re unsure.
- Word count and formatting meet your school’s submission requirements, including page numbering and an authenticity statement.
Two rapid proofreading tactics: Read your evaluation out loud from the conclusion backward. Sentences that sound vague when isolated are the ones examiners flag. Then search the document for the word “error” — every instance needs a specific noun in front of it.
A tutor’s perspective on the one edit that changes everything
What actually separates a mid-band evaluation from a top-band one is almost always a single edit: replacing directional vagueness with a specific, bounded claim about effect. Students write “this may have affected the results.” Examiners want to know how, by how much, and in which direction. That one shift, applied consistently across 2–3 limitations, is worth more than adding extra points.

As an IB examiner with over 13 years of tutoring experience, I see this in nearly every draft I review. The student has the right idea but stops one sentence too early. A targeted session focused on the evaluation section alone — reading each limitation and asking “what does this mean for the RQ?” — can raise a Criterion F score by one or two mark bands in under an hour.
How Esstutor can help you polish your evaluation
Writing a strong ESS evaluation is a skill, and it gets much easier once you’ve seen your own draft marked against the IB criteria by someone who knows exactly what examiners look for.

Esstutor offers targeted IB ESS IA tutoring sessions where Marija, an IB examiner with 13+ years of experience, works through your evaluation paragraph by paragraph. She identifies where you’ve lost the link to your RQ, where your limitations need a direction and effect, and where your improvements need a feasibility note. Sessions are fully online and flexible, so they fit around your school schedule whether you’re in Vienna, Graz, or anywhere else. You can also browse annotated ESS IA examples to see what top-band evaluations look like before your session. Book a trial lesson at Esstutor and bring your current draft — one focused session is often all it takes.
Sources
- Scoring ESS SL’s Data and Evaluation Questions
- Leitfaden Essay Bewertung (GTE) — University of Oldenburg (PDF)
- Einführung in das Schreiben wissenschaftlicher Essays — University of Göttingen (PDF)
No Comments