> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sightlinebehavior.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Writing Up Observation Results

> Turn one session's data into report-ready sentences scoped to what a single observation can support.

<Card title="When to use" icon="circle-question">
  * Turning a completed session's data into report or IEP-ready sentences
  * Deciding what a single session supports versus what needs repeated sessions
  * Choosing which lead statistic to report for a given method
  * Stating within-session patterns as facts instead of function claims
</Card>

Recording a session gets you numbers and notes. Writing it up is a separate skill: turning a count, a set of durations, a run of intervals, or a page of ABC or anecdotal entries into sentences a reader can trust and act on. This guide covers the conventions that apply across methods, then the specific patterns for each one.

## Foundations for any method

### One session tells you what happened, not why

A single count, duration, percentage, or sequence of events describes this observation, under these conditions, on this day. Don't describe a single session's data as increasing, decreasing, or improving, and don't name a function, a pattern that generalizes across settings, or what's typical for the student from one session alone. A comparison to a prior session or a stated goal is fine to report as fact. Conclusions about direction or cause need repeated measures.

### Objective language is a validity issue

Not just a style preference. Record what you observed and tagged, not what you decided it meant. "Escape-motivated" is a conclusion. "The demand was withdrawn after the behavior occurred" is what you saw. Score what you see and hear, and let the write-up's interpretation, if any, sit clearly apart from the description.

### Save interpretation for the evaluation level

A session write-up is declarative: it reports what happened. The interpretive layer of a report, function hypotheses, what's typical for the student, why a behavior occurs, is built at the evaluation level from repeated sessions and multiple data sources, and that's where it belongs. Softening a why-claim doesn't fix its scope: "suggests an escape function" from one session is the same overclaim as "is escape-maintained," in gentler wording. When a session shows a within-session pattern, state it as a count ("a demand preceded 4 of 5 instances"). That's a fact, and it needs no softening.

## Frequency

State the count, the observation window, and the rate together in one sentence, so a reader can check the arithmetic: "Call-outs occurred 14 times during the 20-minute observation, a rate of 0.7 per minute."

Convert to rate when it does work: when you're comparing sessions of different lengths, tracking change against a stated goal, or reporting a figure someone outside this session will need. If every session runs the same fixed length, the raw count already supports comparison, and a rate just adds a division without adding information.

Zero occurrences is a result, not a gap. If the target behavior didn't happen, say so directly: "No instances of call-outs were observed during the 25-minute period." That's a finding a team can act on. Report a peer rate, when one exists, as plain context: state both rates and the difference, and let the reader weigh it.

The weakest, most common pattern is the bare count: "12 tantrums," with no window, no rate, and nothing to compare it to. It can't be checked or compared, including against the student's own other sessions. Name the observation window every time, even when it's the length you always use. See [Frequency Recording](/frequency-recording).

## Duration

Pick one lead metric per behavior, not all three, and let the behavior type decide which:

* **Continuous or engagement behaviors** (on-task, engagement, off-task): lead with **percentage of session**. That's the humanized form of total duration, not a third number stacked on top of it.
* **Episodic behaviors** (tantrums, aggression, crying) that you're also counting: lead with **mean duration per episode**, not total duration. Total duration answers how much of the session the behavior consumed, which is the wrong question when what you're characterizing is how bad a typical episode is.

A percentage of session means nothing without its denominator, and give both sides of the split, not just the side you measured: "off task for 14 of 20 minutes (70%), and on task the remaining 6 minutes (30%)," not just "off task 70% of the session." When a behavior occurred more than once, report the range before the mean, since a mean alone hides whether episodes cluster near it or an outlier is dragging it up. Round to whole minutes once a duration passes about a minute. Keep sub-minute durations in seconds. Keep total duration and duration-per-occurrence in separate clauses so it's clear which number is which.

> **Worked example.** "During the 20-minute independent work period, Marcus was off-task for 14 of the 20 minutes (70%) and on task the remaining 6 minutes (30%). The off-task time accrued across 6 episodes, ranging from 30 seconds to 6 minutes, with most lasting 2 minutes or less. In comparison, a randomly selected peer was off-task for 2 minutes total (10%), across 3 brief episodes, during the same period."

A duration threshold in your operational definition, like a tantrum lasting longer than one minute, doesn't obligate a duration statistic in the writeup either. If the session's most informative story is a frequency count or an interval percentage, report that instead of forcing a duration number to appear just because the definition mentions time.

The weakest pattern here is a bare number with nothing to compare against: "the tantrum lasted 4 minutes," with no session length, no episode count, no criterion. Sightline has the session length and episode count on hand, so use them. See [Duration Recording](/duration-recording).

## Latency

There's little established convention for writing up a single session of latency data, unlike frequency or duration. What follows is a recommended approach for a clear, complete write-up, not a documented field standard.

Tie mean, median, range, and N together in one sentence: "\[Student] took a mean of \[X] seconds to \[target behavior] across \[N] opportunities (median \[Y] seconds, range \[A] to \[B] seconds)." A mean without N is hard to interpret: 22 seconds across 3 opportunities and 22 seconds across 20 carry very different confidence, even though the sentence looks the same.

State how no-response trials were handled. Sightline calculates mean, median, and range from trials with a recorded response only. No-response trials are counted separately and excluded from those statistics. Say so directly: "Two additional opportunities had no response within the observation window and are not included in these figures."

Say which direction is good. A latency number is ambiguous on its own: shorter is better for compliance or task-initiation targets, but for latency to the onset of problem behavior, how long a student tolerates a demand before the behavior begins, longer is better. Name the direction when it isn't obvious from context. Pair the number with a criterion or baseline when one exists, such as a compliance criterion or a peer figure you've gathered separately (peer comparison isn't a built-in Sightline latency feature, so that data is yours to bring).

> **Worked example.** "Malik took a mean of 22 seconds to begin his independent work across 9 opportunities (median 18 seconds, range 6 to 55 seconds). Two additional opportunities had no response within the observation window and are not included in these figures. Per his teacher, students in this classroom typically begin within about 10 seconds of a direction."

Keep the write-up short, typically 2 to 4 sentences: the summary sentence, a note on any no-response trials, and one sentence of direction or context is usually enough. The most common gap in real latency write-ups is a mean with no N stated: "Mean latency was 12 seconds" tells a reader nothing about how many opportunities it's based on. See [Latency Recording](/latency-recording).

## Interval methods

A complete interval writeup introduces the method and its bias once, ties each percentage to the method and the interval length in the same sentence, and keeps peer comparison in proportion.

### Say the interval, not the time

Whole interval and partial interval percentages describe intervals scored, not time elapsed: "on-task in 35% of intervals," never "on-task 35% of the time." Reporting a partial interval percentage as a duration percentage is the single most common interval-reporting error. Momentary time sampling comes closer to a true duration estimate, so "approximately 42% of the observation" is defensible for MTS specifically, but the same looseness overstates what whole and partial interval data actually show.

### Name the interval length every time

State it in the same sentence as the percentage: "Off-task behavior was scored in 65% of 15-second intervals across a 20-minute session." A 60%-of-intervals figure means something different at 10 seconds than at 30. And give each behavior its own paragraph. Each carries its own scoring rule and bias direction.

### State the bias once, where you introduce the method

Every interval percentage is a biased estimate of true occurrence in a known direction: whole interval underestimates, partial interval overestimates (sharply, for brief low-rate behaviors), and momentary time sampling comes closest to a true duration estimate. These figures are estimates, never direct measurements, and the acknowledgment belongs at the top of the observation section, where the report introduces the method: "Behavior was recorded using partial-interval recording in 15-second intervals, a method that tends to overestimate the occurrence of brief behaviors." One sentence there covers every figure that follows. What doesn't work is repeating the caveat beside each percentage in the results, which reads as hedging your own data. Within the results, the discipline is the two rules above: percentage of intervals, with the interval length attached.

Sightline's AI-drafted summaries intentionally leave this disclosure out: the drafting model doesn't know where a particular observation summary will land in your report, so it can't place a once-per-section note correctly. We still recommend including a note to this effect in your report as best practice.

### Frame peer comparison as context, not a verdict

Report both raw percentages, then one plain-language sentence connecting them: "Target on-task 35%, peer on-task 80%: the target student was on-task less than half as often as the peer." Sightline shows the target and peer percentages side by side with the difference between them. Lead the writeup with the raw percentages and a plain comparison sentence rather than a difference or ratio on its own, since a derived figure by itself doesn't show a reader what either number actually was. Peer comparison fits classroom engagement targets well. It fits less naturally in 1:1 or problem-behavior contexts with no comparable peer, so skip it rather than force one.

### Disclose any change in method or interval length

If a later session switches from whole to partial interval, or from 15-second to 30-second intervals, say so directly in the results, the same way you'd disclose a change to an operational definition. See [Interval Recording](/interval-recording).

## ABC and anecdotal

Both methods produce a concrete sequence of events, and both write up as plain narration of that sequence. The difference is what else there is to report: an ABC session yields tag counts you can state as facts alongside the account, while an anecdotal write-up is the account itself.

### ABC

From a single ABC session, you can report the specific sequences you observed and the prevalence of each antecedent or consequence tag across the events you recorded ("X occurred before Y in N of M instances this session"). You can't report function, a trend, or whether the pattern generalizes to other settings from one session. Those need repeated sessions.

> **Worked example.** "A demand to begin independent work preceded 4 of 5 recorded instances of work refusal this session. In each of those instances, an adult withdrew or restated the demand within about 30 seconds of the refusal."

Notice neither sentence says "escape function." They state the pairings plainly, because the pairings are what was observed, and don't reach for a reason one session isn't enough to name.

Sightline's own results screen holds to the same standard: it surfaces a function hypothesis ("may be consistent with escape, attention, tangible, or automatic") only at strong or moderate confidence, shows "no dominant pattern yet" otherwise, and its results footer reads verbatim: hypotheses confirm across sessions, not within one. Treat that as the bar your writeup should meet, too, not a label one session earns on its own.

### Anecdotal

An anecdotal write-up's product is the sequenced, objective account itself: what happened, in what order, and under what conditions, during that observation window. Write it declaratively. "He started writing within 5 seconds of the direction" is a complete, reportable sentence, and a write-up that ends on the observed sequence is finished, not missing its conclusion. Real school-evaluation observation narratives are written exactly this way.

> **Worked example.** "During whole-group instruction, Marcus answered two teacher questions and kept his eyes on the board or the teacher throughout. When independent practice began, he started writing within 5 seconds of the direction. He stopped working three times during the 15-minute work period, each time immediately after erasing an error. Twice he turned to talk to the student seated beside him. Each time the teacher restated the step he had missed, Marcus returned to writing within a few seconds."

Resist the interpretive close. "The pattern suggests difficulty maintaining focus after task errors" is a claim about why, and one session can't back it no matter how gently it's worded. That sentence belongs in the evaluation's synthesis, where repeated sessions and other data sources can support it. If the session showed a pattern worth pointing at, state it as what it is: "disengagement followed task errors in all three instances" is a fact, and it hands the evaluation exactly what it needs.

### Patterns to avoid in both

* Restating the log with no synthesis: a list of tags or entries isn't a summary, it's the data you already have.
* Naming a function or a firm pattern from too few instances: escape, attention, tangible, and automatic are conclusions a repeated pattern across sessions can support, not a label one session earns.
* Manufacturing a pattern that isn't there. "No clear pattern emerged in this session" is a legitimate, useful sentence when the events don't share a clear antecedent or consequence.

See [ABC Recording](/abc-recording) and [Anecdotal Recording](/anecdotal-recording).
