Skip to content

Who acted in a shared AI desktop session?

Our prototype tags every input event as human or agent and enforces human override before input reaches the desktop. On Windows, each action is recorded with the interface element it targeted.

Next, we’re linking actions to the network requests they cause.

One desktop can have more than one actor

A person opens a customer record. An AI agent clicks a button in the same desktop. To the application, both actions come from the same logged-in account.

That helps collaboration, but account identity alone cannot tell the person from the agent. The session itself has to record who acted.

We built that record in layers. Input-source tags, an ordered action journal and Windows interface observations place each action in context. Network capture records what left the hosted browser.

Knowing where an action came from

Every input frame sent to the remote desktop starts with one byte, human or agent. The session proxy reads it before the event reaches the desktop and writes the action, with its actor, to an ordered journal.

Tagging is per event. A person typing while the agent is halfway through a drag is still recorded as the person. In the prototype the client sets the tag, and we’re moving attribution to a trusted server-side path.

The tag says the agent clicked at (512, 340). It does not say that the click submitted a form with a customer’s details. The next two layers record that context.

PersonInput tagged human
AI agentInput tagged agent
Tagged input frames

Session proxyWindows prototype

Reads the source tag
Enforces human overrideBlocked agent input is dropped
Ordered journalEach action with its actor
Allowed input
Desktop
Network captureA separate session-level stream, with no actorOur Chromium build
Each input frame carries a human or agent tag. The session proxy reads it, drops blocked agent input and journals each action with its actor. Network capture is a separate session-level stream with no actor.

The person needs a dependable handover

Human override is enforced at the session proxy, before input reaches the desktop. Blocked agent input is dropped there, so pausing the agent does not depend on the agent noticing and cooperating.

Three modes decide whose input goes through. In Watch, the agent drives. In Co-drive, any human input pauses the agent, which may resume after 1.2 seconds of human inactivity. In Drive, the agent is paused.

Pausing the agent also releases any keys and mouse buttons it was holding. Blocking new input alone would leave a button pressed from an earlier action.

Override and attribution are separate records. Attribution says who acted. Override says whether that actor was allowed to act at that moment. The journal keeps both.

Override modes, enforced at the session proxy
ModeWhose input goes through
WatchThe agent drives
Co-driveAny human input pauses the agent, which may resume after 1.2 seconds of human inactivity
DriveThe agent is paused

A click needs context

Coordinates alone are a poor explanation of an action. The same point on a screen can refer to a different button after a window moves or a page changes.

On Windows, a service in the Virtual Machine sends window information and the UI Automation tree over an RDP virtual channel, NI::CTX, on the same connection as the video. The agent acts on numbered marks tied to a snapshot. A click on a mark from an older snapshot is rejected instead of landing on whatever moved into that spot.

A click on a mark becomes an ordinary click, so it follows the same override and recording path and carries the same source tag. The journal records input and interface context together, in one coordinate space.

Capturing what leaves the browser

For our own Chromium build, a component build that ships libthird_party_boringssl.so with symbols, we built an eBPF tap that reads TLS traffic in plaintext, with no proxy and no man-in-the-middle certificate. Uprobes on Chromium’s BoringSSL copy each read and write into a ring buffer. The analyzer stores HTTP/1, HTTP/2, WebSocket and SSE request and response pairs.

The tap has no view of who caused a request. By the time traffic reaches BoringSSL, Chrome’s network service has merged every tab and both actors into one stream. Network records are kept as a separate session-level stream.

Linking an action to what it caused

Suppose a customer record is sent to an external service during a shared session. The investigation needs to know whether an agent action caused that transfer.

Applications make background requests, delay work and react to several events at once, so a request close in time to a click can be unrelated. Timing alone is weak evidence of cause.

We are working through three approaches. One attributes a request that starts within a set window of a tagged click in the same tab. Another uses Chrome’s own initiator chain over the DevTools Protocol. Where the target site allows it, OAuth on-behalf-of carries the agent’s identity in the request itself.

Reading a session as workflows

A longer sequence of actions often follows a familiar workflow. Recognizing it lets a reviewer read a session as workflows rather than event by event.

Our design finds workflows by compression. Each event becomes a role-and-verb token, so two different customer record URLs map to one token. A repeated sequence becomes a workflow only when it compresses the stream. Rare workflows never compress well, so people can confirm rules directly.

An inferred workflow never grants permission. It helps a reviewer read the record. Whether the agent may act is decided separately, at the session proxy.

What we’re building next

The Windows prototype records per-event source tags, enforced human override and interface context in one journal. Network capture runs in our own Chromium build.

Next, we’re building the link from an action to the requests it caused, interface capture for the hosted browser and the workflow model.

The evaluation that follows measures attribution accuracy, memory use and performance overhead.

All research