> ## Documentation Index
> Fetch the complete documentation index at: https://docs.glasswarp.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Observe the screen

> Capture frames, regions of interest, text summaries, and only what changed.

## One-shot screenshot

```python theme={null}
frame = gw.screenshot(sid)
with open("screen.jpg", "wb") as f:
    f.write(frame.jpeg)
```

Downscale for cheaper model calls, or crop to a region of interest:

```python theme={null}
small = gw.screenshot(sid, max_width=1280, quality=80)
roi = gw.screenshot(sid, x=100, y=200, w=800, h=600)
```

## observe — frame + change + targets + text

```python theme={null}
obs = gw.observe(sid, max_width=1280, mark=True)
print(obs.text.summary if obs.text else "")
print(obs.changed)          # False only when DXGI dirty available + empty
print(obs.capture_mode)     # "dxgi" (preferred) or "gdi_fallback"
print(obs.dirty)            # None when dirty unavailable — assume changed
print(len(obs.jpeg), "bytes")
print(len(obs.targets), "targets")
```

`obs.text` is a compact structured summary — focused window title/role plus
readable target lines like `[3] Filename (edit, focused)` — so your model can
often decide without spending image tokens.

Password and masked fields are redacted in that text (`masked`, value always
`[redacted]`). See [Safety and consent](/concepts/safety-and-consent).

Pass your own targets to get a Set-of-Mark annotated frame back:

```python theme={null}
from glasswarp import targets_from_grid

targets = targets_from_grid(x0=100, y0=200, x1=900, y1=800, rows=8, cols=8)
obs = gw.observe(sid, mark=True, targets=targets)
```

## Text-only verification (`image=False`)

For simple checks — did the dialog close? is the field focused? — skip the JPEG:

```python theme={null}
obs = gw.observe(sid, image=False)
print(obs.text.summary if obs.text else "")
print(obs.changed)
# obs.jpeg is empty
```

This is much faster than a full frame. Request the image when you need to read
or judge the screen visually.

## Cheap verification frames

When you do need a JPEG for a quick verify, downscale harder:

```python theme={null}
obs = gw.observe(sid, max_width=960, quality=60, mark=True)
```

## Only look when something changed

```python theme={null}
rects = gw.dirty_rects(sid)
if rects:
    # Encode only the changed region when it is small:
    x, y, w, h = rects[0]
    frame = gw.screenshot(sid, x=x, y=y, w=w, h=h, max_width=960, quality=60)
else:
    ...  # nothing changed; skip the expensive call
```

Or use `observe` and check `obs.changed` / empty `obs.dirty["rects"]`.

<Tip>
  Pair dirty rects with ROI observe (`x,y,w,h`): encode only the changed region
  and skip full-frame model calls.
</Tip>

## Cadence

You do not need a fresh observe between every click. Verify after each
**meaningful step** — after a short batch of predictable actions, or after a
single action when the next screen is uncertain. Prefer `image=False` for
verification when text/targets are enough.

Target ids from `observe` / `list_targets` are valid for that response only. Do not
persist them across sessions, and re-observe before clicking when the screen may
have changed (lists, file browsers, tables).

See [Vision](/concepts/vision) for the concepts and [Agent
loop](/guides/agent-loop) to put it together.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.