All updates
2026.09.27

Agent Work You Can Measure, Sessions That Don't Get Stuck

Evaluations turn agent output into a graded workflow, a reliability overhaul keeps long sessions alive through stalls and restarts, and desktop notifications now take you straight to the conversation.

Features Improvements Fixes

Grade the work, not just the words

Evaluations grade agent sessions against a rubric

Evaluations grew into a real workflow this cycle. You can now create them at the organization level with a setup wizard, and the agent itself runs the evaluation through a dedicated tool surface — replaying its own sessions against your rubric and writing the results to your dashboard, where a human makes the final call. Orchestration project sessions can be evaluated too, so a whole multi-agent run gets graded as one piece of work. If your agent runs on compatible subagent infrastructure, that carries over as well.

The upshot: “it looks right” stops being the only check you have. You get a score, the specific rubric lines that failed, and a queue of sessions worth a second look.

Sessions that pick themselves back up

Dropped turns resume automatically on a session timeline

Long-running sessions used to have a quiet failure mode: a backend restart or a stalled stream would kill a turn mid-flight, and you’d find out from a missing result. This cycle ships a reliability wave across the whole path. Dropped turns now auto-continue instead of dying. Turns left dangling by a restart get reaped when the session goes idle again. An enqueue deadlock that could wedge scheduling is gone, and completion notices arrive exactly once — never doubled, never lost. The watchdog got honest about stream stalls, naming what it killed and why, and plan approval cards survive the restarts that used to break them.

Notifications that take you there

A desktop notification that jumps into the exact session

Desktop alerts stop being dead-end banners. Click one and you land in the exact session it came from. Alerts are named after their session, plan approvals show up as alerts with the context you need, and macOS agent updates now install silently in the background instead of interrupting you.

Nothing gets thrown away

Archived sessions are now browsable and restorable from the UI, so an old run you killed can come back without rebuilding it. They stay out of orchestration lists, and the fork menu collapses into sections to keep long histories readable.

Schedules, one honest selector

The schedule target selector merged into a single control with an Auto default that falls back in a predictable order, runtime and permission settings live in one place, and the timezone handling does what it claims.

Also shipped

  • Mobile, one continuous boot — startup is a single flow now, with splash fixes for the three-stage startup defect
  • DeepSeek (dsh) approval flows — approvals and context usage work correctly on the DeepSeek runtime
  • Restore where you were — the app reopens to the last page you were on after a restart
  • Org-owned credentials — credentials now belong to the organization, shared across agents
  • Plan mode for Cursor and Codex — the same planning gate those runtimes get
  • Six releases (0.127.0 → 0.133.0) — including a fix that lets a bad frontend bundle be rolled back without a full redeploy

Start with CrossMind.

Start for Free