> ## Documentation Index
> Fetch the complete documentation index at: https://docs.glasswarp.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Build an agent loop

> observe → think → act → verify after each meaningful step.

Every Glasswarp integration is the same shape. Glasswarp owns **observe** and
**act**; your model owns **think** and **verify**. Open-loop scripts (hardcoded
coords, “always click 400,300”) break the moment scale or layout changes — fine
as helpers *after* vision/UIA chose a region, not as the outer brain.

## The shape

```
Observe  — screenshot / dirty / ROI / UIA targets
Think    — where is the thing? is it safe to click? did the last step work?
Act      — one action, or a short batch when the next screens are predictable
Verify   — look again after each meaningful step; correct or stop
```

Verify after each **meaningful step** — after a batch of predictable actions, or
after a single action when the screen may change unpredictably (page loads,
installers, network waits, anything that can pop a modal). You do not need a
fresh observe between every click in a menu path or dialog you already planned.

```python theme={null}
import os
from glasswarp import GlasswarpClient

gw = GlasswarpClient(api_key=os.environ["GLASSWARP_API_KEY"])

rig = next(
    (r for r in gw.list_rigs() if r.online and r.api_access_enabled),
    None,
)
if rig is None:
    raise RuntimeError("No online rig with API access enabled")

session = gw.create_session(rig_id=rig.id, mode="desktop")
sid = session.session_id

try:
    for step in range(50):
        obs = gw.observe(sid, max_width=1280, mark=True)   # eyes
        action = decide(obs)                               # your brain
        if action is None:
            break
        act(gw, sid, action)                               # hands (one or a short batch)
        # verify after this meaningful step before the next decide
finally:
    gw.end_session(sid)                                    # always clean up
```

`decide` and `act` are yours: `decide` calls your model with `obs.jpeg` and
`obs.targets`; `act` maps its output to `click_target`, `type_text`, or a
batched `send_input` / MCP `send_actions` when the next few steps are known.
After a meaningful step, **look again** — dialogs move, folders change, and
“I clicked Desktop” is not the same as “Save As is on Desktop.”

## Batch when you can predict

If the next 3–6 actions are predictable (menu paths, dialog fields, typing into
a field you just focused), send them as one batch, then observe once to verify
the whole sequence. Single-step when intermediate state is uncertain. If the
verifying observe shows something unexpected, re-plan from there.

MCP clients use `send_actions` (verification observe on by default — text and
targets; a JPEG only when you set `observe_image=true`). SDK clients use
`send_input` with an ordered `events` list — same host path.

## Make it cheap: skip idle frames

```python theme={null}
for step in range(200):
    if not gw.dirty_rects(sid):
        continue                       # nothing changed — no model call
    obs = gw.observe(sid, mark=True)
    ...
```

<Tip>
  The example loops skip idle frames by default. Empty `dirty_rects` means
  "don't spend a model call."
</Tip>

## Ground every action

Prefer `click_target` over raw pixels so the loop survives layout shifts:

```python theme={null}
obs = gw.observe(sid, mark=True)
target_id = decide(obs)                 # model returns a mark id from *this* observe
gw.click_target(sid, target_id, targets=obs.targets)
```

Target ids are valid for the observe that produced them. Do not persist them across
sessions, and re-observe before acting when lists, file browsers, or tables may have
changed. Scaffolded agents must re-resolve targets each loop — never hardcode ids.

## Keep task logic out of the transport

<Note>
  Put prompts, CV, and planning in your own module. The Glasswarp client is a
  thin transport adapter — mixing task logic into it makes both harder to test.
</Note>

## Reference loops

The SDK ships runnable agent loops:

* `templates/observe_think_act/` — bare agent-loop skeleton
* `minesweeper_solver_demo.py` — demo with CV + deterministic solver
* `gemini_agent_loop.py` — demo with Gemini + SoM → `click_target`
* `claude_agent_loop.py` — observe → Claude → act (skips idle frames by default)
* `paint_mona_lisa_demo.py` — full demo on a real app (place → paint → Save As)

See [Templates](/examples/templates), [Minesweeper](/examples/minesweeper), and
[Mona Lisa](/examples/mona-lisa).

Before you turn autonomy up, read [Safe agent
sessions](/guides/safe-agent-sessions) — budgets, human-in-the-loop review, and allowlists on top of
platform consent and Live View.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.