---
title: "Agent TARS & UI-TARS Desktop Monorepo"
description: "An implementation specification for rebuilding a multimodal AI agent monorepo — Agent TARS (CLI + Web UI) and UI-TARS Desktop (Electron GUI automation)."
---

> Documentation Index
> Fetch the complete documentation index at: https://papers.ahmad.wtf/llms.txt
> Use this file to discover all available pages before exploring further.

# Agent TARS & UI-TARS Desktop Monorepo

> **Mode: Implementation Specification** — Instructions for an AI coding agent to rebuild this from scratch. Not a summary or description.

## 1. Product Goal

Build a multimodal AI Agent monorepo that ships two primary products: (1) **Agent TARS** — a CLI and Web UI that brings GUI Agent and Vision capabilities into terminal, browser, and product workflows via cutting-edge multimodal LLMs and MCP tool integration; and (2) **UI-TARS Desktop** — a native Electron desktop application providing GUI automation driven by the UI-TARS Vision-Language Model to control local and remote computers and browsers. Target users are developers and automation engineers who need human-like task completion through natural language instructions. The system must support multiple model providers (OpenAI-compatible, Anthropic, Volcengine, HuggingFace), hybrid browser control (GUI Agent, DOM, or hybrid strategy), and seamless MCP server mounting.

Implement the following core features:

- Support one-click CLI execution with both headful Web UI and headless server modes
- Implement a hybrid browser agent controlling browsers via GUI grounding, DOM extraction, or hybrid strategy
- Build an Event Stream protocol driving context engineering and an Agent UI
- Support mounting arbitrary MCP Servers (stdio, SSE, streamable-HTTP) to extend agent capabilities
- Implement local and remote computer operators (ADB, nut-js, browser-operator, browserbase)
- Create an SDK abstraction layer for GUI agents with pluggable model and operator implementations
- Build a visualizer component for replaying and inspecting agent execution dumps
- Support OpenAI Chat Completions API and Response API for model invocation with sliding image context management

## 2. Tech Stack & Versions

- **Node.js** ≥22.x — primary runtime (CLI requires ≥22, monorepo engines ≥20)
- **TypeScript** ^5.7.2 — language across all packages
- **pnpm** 9.10.0 — workspace package manager
- **Turborepo** ^2.4.4 — monorepo build orchestration
- **Rslib** (`@rslib/core`) — library bundling for all internal packages (ESM + CJS + DTS)
- **Vitest** ^3.0.x — unit, integration, and benchmark testing
- **Electron** — desktop application runtime (UI-TARS Desktop)
- **React** ^18.x — UI renderer (Visualizer, Agent UI)
- **Ant Design** ^5.x — component library for Visualizer
- **Zod** — runtime schema validation across MCP tools, agent actions, and IPC
- **OpenAI SDK** (`openai`) — model provider client (Chat Completions + Responses API)
- **Puppeteer-core** — browser automation engine
- **@modelcontextprotocol/sdk** — MCP client/server protocol implementation
- **ESLint** ^8.57.0 + **Prettier** ^3.3.2 — code quality
- **Husky** ^9.1.7 + **lint-staged** ^14.0.1 — git hooks
- **Changesets** ^2.27.11 — versioning and publishing
- **Playwright** ^1.49.1 — E2E testing for Electron app

## 3. Project Setup & Commands

Initialize the monorepo with the following structure and commands:

```bash
# Create root workspace
mkdir monorepo && cd monorepo
git init
pnpm init

# Install core dev dependencies
pnpm add -Dw turbo@^2.4.4 typescript@^5.7.2 @types/node@^20.14.8
pnpm add -Dw eslint@^8.57.0 @typescript-eslint/eslint-plugin@^5.0.0 @typescript-eslint/parser@^5.0.0 eslint-plugin-import@^2.25.0 eslint-plugin-react@^7.34.3
pnpm add -Dw prettier@^3.3.2 @trivago/prettier-plugin-sort-imports@^5.2.1
pnpm add -Dw vitest@^3.0.8 @vitest/coverage-v8@^3.0.8
pnpm add -Dw husky@^9.1.7 lint-staged@^14.0.1 @commitlint/cli@^19.6.1 @commitlint/config-conventional@^19.6.0
pnpm add -Dw @changesets/cli@^2.27.11 secretlint@^10.2.1
pnpm add -Dw tsx@^4.19.2 ts-node@^10.9.2 rimraf@^6.0.1 cross-env@^7.0.3 sass-embedded@^1.83.1
pnpm add -Dw @playwright/test@^1.49.1 electron-playwright-helpers@^1.7.1
pnpm add -Dw opencommit@^3.2.5 @electron-toolkit/tsconfig@^1.0.1
```

Define the following `package.json` scripts at the root:

```json
{
  "scripts": {
"bootstrap": "pnpm i",
"dev:ui-tars": "turbo run ui-tars-desktop#dev",
"format": "prettier --write .",
"lint": "eslint . --ext .js,.jsx,.cjs,.mjs,.ts,.tsx,.cts,.mts --fix",
"test": "vitest",
"test:bench": "vitest bench",
"coverage": "vitest run --coverage",
"prepare": "husky"
  }
}
```

Create `pnpm-workspace.yaml`:

```yaml
packages:
  - 'apps/*'
  - 'apps/agent-tars/src/*'
  - 'apps/ui-tars/src/*'
  - 'packages/ui-tars/**'
  - 'packages/agent-infra/**'
  - '!packages/agent-infra/create-new-mcp/template-*'
  - 'packages/common/*'
```

Create the following directory structure:

```
apps/
  agent-tars/          # Agent TARS application (CLI + Web UI + Server)
  ui-tars/             # UI-TARS Desktop Electron app
packages/
  agent-infra/         # Agent infrastructure libraries
browser/           # Browser automation (LocalBrowser, RemoteBrowser)
browser-use/       # Browser agent with DOM extraction
mcp-servers/       # Built-in MCP servers (filesystem, etc.)
mcp-client/        # MCP client implementation
mcp-http-server/   # MCP HTTP server launcher
mcp-shared/        # Shared MCP types
logger/            # Logging infrastructure
search/            # Search providers (Bing, DuckDuckGo, browser-search)
operators/         # Computer operators (adb, nut-js, browser-operator, browserbase)
shared/            # Shared agent-infra utilities
create-new-mcp/    # CLI scaffolding tool for new MCP servers
  ui-tars/
sdk/               # GUIAgent SDK (Model, Operator abstractions)
action-parser/     # Parse VLM predictions into actions
visualizer/        # React-based execution dump visualizer
cli/               # UI-TARS CLI
electron-ipc/      # Type-safe Electron IPC
utio/              # Telemetry/analytics
  common/
electron-build/    # Electron build utilities
configs/           # Shared build/lint configs
```

## 4. Environment Variables

```bash
VLM_PROVIDER=huggingface                    # Model provider identifier — required
VLM_BASE_URL=http://endpoint/v1             # OpenAI-compatible base URL — required
VLM_API_KEY=hf_xxx                          # API key for model provider — required
VLM_MODEL_NAME=your_model_name              # Model name to invoke — required
OPENAI_API_KEY=sk-xxx                       # OpenAI API key (when provider=openai) — optional
ANTHROPIC_API_KEY=sk-ant-xxx                # Anthropic API key (when provider=anthropic) — optional
VOLCENGINE_API_KEY=xxx                      # Volcengine API key (when provider=volcengine) — optional
DEBUG=mcp                                   # Enable debug logging for MCP client — optional
UTIO_ENDPOINT=https://utio.example.com/api  # Telemetry endpoint — optional
```

## 5. Key Dependencies

**Model & AI:**

- `openai` — OpenAI-compatible client for Chat Completions and Responses API
- `zod` — Schema validation for all tool inputs, agent outputs, IPC contracts

**Browser Automation:**

- `puppeteer-core` — Headless/headful browser control via CDP
- `@agent-infra/browser` (workspace) — LocalBrowser, RemoteBrowser, BrowserFinder abstractions

**MCP Protocol:**

- `@modelcontextprotocol/sdk` — Client (stdio, SSE, streamable-HTTP transports), Server, InMemory transport
- `uuid` — Generate unique tool IDs
- `minimatch` — Glob pattern matching for MCP tool filtering

**CLI Interaction:**

- `@clack/prompts` — Interactive CLI prompts for create-new-mcp scaffolding
- `mri` — CLI argument parsing
- `picocolors` — Terminal color output

**Electron Desktop:**

- `electron` — Desktop runtime
- `@electron-toolkit/tsconfig` — Shared TS configs for Electron

**UI/Visualizer:**

- `antd` — Component library (Modal, Input, Button, Upload, ConfigProvider)
- `react-resizable-panels` — Resizable panel layouts in Visualizer
- `@midscene/core` — Execution dump types for Visualizer compatibility

**Build Tooling:**

- `@rslib/core` — Library build system producing ESM + CJS + DTS bundles

## 6. Data Models & Schema

### MCP Filesystem Server Schemas (`packages/agent-infra/mcp-servers/filesystem/src/schema.ts`)

```typescript
import { z } from 'zod';

export const ReadFileArgsSchema = z.object({ path: z.string() });
export const ReadMultipleFilesArgsSchema = z.object({ paths: z.array(z.string()) });
export const WriteFileArgsSchema = z.object({ path: z.string(), content: z.string() });
export const EditOperation = z.object({
  oldText: z.string().describe('Text to search for - must match exactly'),
  newText: z.string().describe('Text to replace with'),
});
export const EditFileArgsSchema = z.object({
  path: z.string(),
  edits: z.array(EditOperation),
  dryRun: z.boolean().default(false),
});
export const CreateDirectoryArgsSchema = z.object({ path: z.string() });
export const ListDirectoryArgsSchema = z.object({ path: z.string() });
export const DirectoryTreeArgsSchema = z.object({ path: z.string() });
export const MoveFileArgsSchema = z.object({ source: z.string(), destination: z.string() });
export const SearchFilesArgsSchema = z.object({
  path: z.string(),
  pattern: z.string(),
  excludePatterns: z.array(z.string()).optional().default([]),
});
export const GetFileInfoArgsSchema = z.object({ path: z.string() });
```

### Browser Agent Action Schemas (`packages/agent-infra/browser-use/src/agent/actions/schemas.ts`)

```typescript
import { z } from 'zod';

export interface ActionSchema {
  name: string;
  description: string;
  schema: z.ZodType;
}

export const doneActionSchema: ActionSchema = {
  name: 'done', description: 'Complete task',
  schema: z.object({ text: z.string() }),
};

export const searchGoogleActionSchema: ActionSchema = {
  name: 'search_google', description: 'Search Google in the current tab',
  schema: z.object({ query: z.string() }),
};

export const goToUrlActionSchema: ActionSchema = {
  name: 'go_to_url', description: 'Navigate to URL in the current tab',
  schema: z.object({ url: z.string() }),
};

export const goBackActionSchema: ActionSchema = {
  name: 'go_back', description: 'Go back to the previous page',
  schema: z.object({}),
};

export const clickElementActionSchema: ActionSchema = {
  name: 'click_element', description: 'Click element',
  schema: z.object({
desc: z.string().optional(),
index: z.number(),
xpath: z.string().optional(),
  }),
};

export const inputTextActionSchema: ActionSchema = {
  name: 'input_text', description: 'Input text into an interactive input element',
  schema: z.object({
desc: z.string().optional(),
index: z.number(),
text: z.string(),
xpath: z.string().optional(),
  }),
};

export const switchTabActionSchema: ActionSchema = {
  name: 'switch_tab', description: 'Switch to tab by id',
  schema: z.object({ tab_id: z.number() }),
};

export const openTabActionSchema: ActionSchema = {
  name: 'open_tab', description: 'Open URL in new tab',
  schema: z.object({ url: z.string() }),
};

export const scrollDownActionSchema: ActionSchema = {
  name: 'scroll_down',
  description: 'Scroll down the page by pixel amount',
  schema: z.object({ desc: z.string().optional(), amount: z.number().optional() }),
};

export const scrollUpActionSchema: ActionSchema = {
  name: 'scroll_up',
  description: 'Scroll up the page by pixel amount',
  schema: z.object({ desc: z.string().optional(), amount: z.number().optional() }),
};

export const sendKeysActionSchema: ActionSchema = {
  name: 'send_keys',
  description: 'Send strings of special keys like Backspace, Enter, Control+o',
  schema: z.object({ desc: z.string().optional(), keys: z.string() }),
};

export const scrollToTextActionSchema: ActionSchema = {
  name: 'scroll_to_text',
  description: 'Scroll to specific text on the page',
  schema: z.object({ desc: z.string().optional(), text: z.string() }),
};

export const getDropdownOptionsActionSchema: ActionSchema = {
  name: 'get_dropdown_options', description: 'Get all options from a native dropdown',
  schema: z.object({ index: z.number() }),
};

export const selectDropdownOptionActionSchema: ActionSchema = {
  name: 'select_dropdown_option',
  description: 'Select dropdown option for interactive element index',
  schema: z.object({ index: z.number(), text: z.string() }),
};
```

### Agent Types (`packages/agent-infra/browser-use/src/agent/types.ts`)

```typescript
export interface AgentOptions {
  maxSteps: number;
  maxActionsPerStep: number;
  maxFailures: number;
  retryDelay: number;
  maxInputTokens: number;
  maxErrorLength: number;
  useVision: boolean;
  useVisionForPlanner: boolean;
  validateOutput: boolean;
  includeAttributes: string[];
  planningInterval: number;
}

export const DEFAULT_AGENT_OPTIONS: AgentOptions = {
  maxSteps: 100, maxActionsPerStep: 10, maxFailures: 3, retryDelay: 10,
  maxInputTokens: 128000, maxErrorLength: 400, useVision: false,
  useVisionForPlanner: false, validateOutput: true,
  includeAttributes: ['title','type','name','role','tabindex','aria-label','placeholder','value','alt','aria-expanded'],
  planningInterval: 3,
};
```

### Browser Context Types (`packages/agent-infra/browser-use/src/browser/types.ts`)

```typescript
export interface BrowserContextConfig {
  minimumWaitPageLoadTime: number;     // default: 1
  waitForNetworkIdlePageLoadTime: number; // default: 1.0
  maximumWaitPageLoadTime: number;     // default: 5.0
  waitBetweenActions: number;          // default: 1.0
  browserWindowSize: { width: number; height: number }; // default: 1280x1100
  highlightElements: boolean;          // default: true
  viewportExpansion: number;           // default: 500
  includeDynamicAttributes: boolean;   // default: true
  homePageUrl: string;                 // default: 'https://www.google.com'
}

export interface PageState extends DOMState {
  tabId: number;
  url: string;
  title: string;
  screenshot: string | null;
  pixelsAbove: number;
  pixelsBelow: number;
}

export interface BrowserState extends PageState {
  pages?: PageInfo[];
}
```

### UI-TARS SDK Core Types (`packages/ui-tars/sdk/src/types.ts`)

```typescript
export interface InvokeParams {
  conversations: Message[];
  images: string[];
  screenContext: { width: number; height: number };
  scaleFactor?: number;
  uiTarsVersion?: UITarsModelVersion;
  headers?: Record<string, string>;
  previousResponseId?: string;
}

export interface InvokeOutput {
  prediction: string;
  parsedPredictions: PredictionParsed[];
  costTime?: number;
  costTokens?: number;
  responseId?: string;
}

export interface GUIAgentConfig<TOperator> {
  operator: TOperator;
  model: InstanceType<typeof UITarsModel> | ConstructorParameters<typeof UITarsModel>[0];
  systemPrompt?: string;
  signal?: AbortSignal;
  onData?: (params: { data: GUIAgentData }) => void;
  onError?: (params: { data: GUIAgentData; error: GUIAgentError }) => void;
  logger?: Logger;
  retry?: { model?: RetryConfig; screenshot?: RetryConfig; execute?: RetryConfig; };
  maxLoopCount?: number;   // default: 25
  loopIntervalInMs?: number; // default: 0
  uiTarsVersion?: UITarsModelVersion;
}
```

### MCP Client Types

```typescript
export interface MCPClientOptions {
  isDebug?: boolean;
  defaultTimeout?: number; // seconds, default: 60
}

export interface MCPTool extends Tool {
  id: string;
  serverName: string;
}
```

### UTIO Event Types (`packages/ui-tars/utio/src/types.ts`)

```typescript
export type EventType = 'appLaunched' | 'sendInstruction' | 'shareReport';

export interface AppLaunchedEvent extends BaseEvent<'appLaunched'> {
  platform: string;
  osVersion: string;
  screenWidth: number;
  screenHeight: number;
}

export interface SendInstructionEvent extends BaseEvent<'sendInstruction'> {
  instruction: string;
}

export interface ShareReportEvent extends BaseEvent<'shareReport'> {
  lastScreenshot?: string;
  report?: string;
  instruction: string;
}
```

### Electron IPC Types

```typescript
export type ZodSchema<TInput> = { parse: (input: any) => TInput };
export type HandleFunction<TInput = any, TResult = any> = (args: {
  context: HandleContext;
  input: TInput;
}) => Promise<TResult>;
export type HandleContext = { sender: WebContents | null };
export type RouterType = Record<string, { handle: HandleFunction }>;
```

### UITarsModelVersion Enum

```typescript
export enum UITarsModelVersion {
  V1_0 = 'V1_0',
  V1_5 = 'V1_5',
  DOUBAO_1_5_15B = 'DOUBAO_1_5_15B',
  DOUBAO_1_5_20B = 'DOUBAO_1_5_20B',
}
```

## 7. Database & Storage Setup

N/A — This project does not use a traditional database. State is managed in-memory during agent execution, persisted to local files (execution dumps as JSON), and stored in browser localStorage for the Visualizer environment configuration. File-based storage is used for MCP server configurations and execution reports written to `./midscene_run/report/` directories.

## 8. API Endpoints & Contracts

### MCP HTTP Server Endpoints

```
POST /mcp
Auth: None (configure at deployment)
Request: JSON-RPC 2.0 envelope { jsonrpc: "2.0", method: string, params: object, id: number }
Response: JSON-RPC 2.0 result { jsonrpc: "2.0", result: object, id: number }
Methods: tools/list, tools/call, resources/list, prompts/list, initialize
Transport: Streamable HTTP or SSE
```

### Agent TARS Server Endpoints **[INFERRED]**

```
POST /api/sessions
Auth: None
Request: { instruction: string, provider: string, model: string, apiKey: string }
Response: { sessionId: string }
Errors: 400 when instruction is empty

GET /api/sessions/:id/events
Auth: None
Response: SSE stream of AgentEvent objects
Events: { actor: string, state: string, details: string, browserState?: BrowserState }

POST /api/sessions/:id/stop
Auth: None
Response: { status: "stopped" }
```

### UTIO Telemetry

```
POST {UTIO_ENDPOINT}
Auth: None
Request: EventPayload<T> — { type: EventType, ...eventSpecificFields }
Response: 200 OK (silent fail on error)
```

## 9. Authentication & Permissions

N/A for traditional user auth. This system uses API key-based authentication for model providers:

1. Accept API keys via CLI flags (`--apiKey`), environment variables (`VLM_API_KEY`, `OPENAI_API_KEY`, etc.), or runtime configuration UI
2. Pass keys to the OpenAI client constructor as `apiKey` in `ClientOptions`
3. Never log or expose API keys — mask values as `***` in all UI displays (except `MIDSCENE_MODEL_NAME`)
4. Store environment configuration in browser `localStorage` for the Visualizer component
5. For MCP servers using stdio transport, inject environment variables into the child process with enhanced PATH resolution
6. For HTTP-based MCP servers, pass custom headers via `requestInit.headers`

## 10. Pages & Routes

### Agent TARS Web UI **[INFERRED]**

- `/` — Main chat/agent interface; user inputs instruction, views streaming agent execution, tool calls, and browser screenshots
- `/settings` — Model provider configuration, MCP server management, operator selection
- `/events` — Event Stream Viewer for debugging data flow between agent steps

### UI-TARS Desktop (Electron Windows)

- **Main Window** — Agent UI with instruction input, live screenshot preview, action timeline, and status indicators
- **Report Window** — Visualizer rendering execution dumps with Sidebar, Timeline, DetailPanel, DetailSide, and Player components

### Visualizer Routes (Standalone HTML)

- `/playground.html` — Interactive playground for testing agent flows
- `/report.html` — Static report viewer loading `.web-dump.json` files via drag-and-drop Upload component

## 11. Component Architecture

### Visualizer Component Tree

```
Visualizer (root)
├── ConfigProvider (antd theme: globalThemeConfig)
├── EnvConfig
│   ├── Button (Edit) → Modal with Input.TextArea for KEY=VALUE pairs
│   └── Tooltip for status indicators
├── PanelGroup (main-page-layout, horizontal)
│   ├── Panel (sidebar, 20%)
│   │   └── Sidebar — list of execution tasks
│   ├── PanelResizeHandle
│   └── Panel (main-right, 80%)
│       ├── Timeline — temporal action sequence
│       └── PanelGroup (page-detail-layout-v2)
│           ├── Panel (detail-panel, 75%)
│           │   └── DetailPanel — screenshots, actions, results
│           ├── PanelResizeHandle
│           └── Panel (detail-side)
│               └── DetailSide — metadata and side info
├── GlobalHoverPreview — floating element preview on hover
└── Player — replay mode for execution animations
```

**EnvConfig Props:** None (uses `useEnvConfig` Zustand store)

**EnvConfig State:** `isModalOpen: boolean`, `tempConfigString: string`, `showEditButton: boolean`

**EnvConfig Behaviors:** Load config from textarea, persist to store, mask API key values

**Visualizer Props:** `{ logoAction?: () => void; dumps?: EnhancedGroupedActionDump[]; onActiveTaskChange?: (taskIndex, task) => void }`

### Agent TARS CLI Architecture

```
@agent-tars/cli
├── CLI Parser (commander/yargs)
│   ├── --provider (openai|anthropic|volcengine|huggingface|...)
│   ├── --model
│   ├── --apiKey
│   └── --headless (flag for server mode)
├── AgentKernel
│   ├── MCPClient (mounts tools from MCP servers)
│   ├── BrowserAgent (browser-use or external operator)
│   ├── ModelProvider (OpenAI-compatible)
│   └── EventStream (protocol-driven event emission)
└── Output Renderer
├── Web UI (Express/Hono server + SSE)
└── Terminal (streaming console output)
```

### UI-TARS SDK Architecture

```
GUIAgent (orchestrator)
├── Model (abstract)
│   └── UITarsModel extends Model
│       ├── invokeModelProvider() — OpenAI Chat Completions or Responses API
│       ├── preprocessResizeImage() — compress to maxPixels by version
│       ├── convertToOpenAIMessages() — map conversations + images
│       └── actionParser() — parse VLM prediction into actions
├── Operator (abstract)
│   ├── screenshot() → ScreenshotOutput
│   └── execute(params: ExecuteParams) → ExecuteOutput
├── Loop Controller
│   ├── maxLoopCount (default: 25)
│   └── loopIntervalInMs (default: 0)
└── Event Emitter (onData, onError callbacks)
```

## 12. State Management

- **Visualizer**: Use Zustand for the `useExecutionDump` store managing: `dump`, `_executionDumpLoadId`, `replayAllMode`, `allExecutionAnimation`, `insightWidth`, `insightHeight`, `onActiveTaskChange`. Actions: `setGroupedDump`, `setReplayAllMode`, `reset`, `setOnActiveTaskChange`.
- **EnvConfig**: Use Zustand `useEnvConfig` store managing: `config` (`Record<string, string>`), `configString` (string). Actions: `loadConfig(configString)`.
- **MCP Client**: Internal `Map<string, any>` store for server registry; `EventEmitter` for status changes; `activeServers: Map<ServerNames, {client, server}>` for connected clients.
- **Agent Context**: Class-based mutable state in `AgentContext` holding `taskId`, `browserContext`, `messageManager`, `eventManager`, `paused`, `stopped`, `consecutiveFailures`, `nSteps`, `actionResults`.
- **UITarsModel**: Instance state `headImageContext: { messageIndex: number; responseIds: string[] } | null` for sliding window image management in Responses API mode.
- **CLI/Server**: Use EventEmitter-based Event Stream for real-time streaming to Web UI via SSE.

## 13. UI/UX & Design System

- **Component Library**: Ant Design v5 with `ConfigProvider` wrapping all Visualizer components using `globalThemeConfig`
- **Layout**: `react-resizable-panels` for resizable PanelGroup layouts with `autoSaveId` for persistence
- **Colors**: Status indicators using `iconForStatus('success')` (green check) and `iconForStatus('failed')` (red X); standard antd color palette
- **Typography**: Default antd typography; monospace for config text areas and code display
- **Upload UX**: `Dragger` component accepting `.web-dump.json` files with FileReader parsing and error messaging via `message.error()`
- **Responsive**: Panel-based responsive layout with resize handles; window resize listener with 300ms throttle
- **Animations**: Replay mode with `Player` component animating execution scripts; panel drag transitions
- **Styling**: LESS files (`index.less`) for custom styles; `whiteSpace: nowrap` with `wordWrap: break-word` for config text areas
- **Dark/Light**: Configure via antd `ConfigProvider` theme tokens **[INFERRED]**

## 14. Integrations & External Services

### Model Providers (OpenAI-compatible)

- Install `openai` SDK
- Initialize on server/main process: `new OpenAI({ baseURL, apiKey, maxRetries: 0 })`
- Implement two invocation paths:
1. **Chat Completions API**: `openai.chat.completions.create({ model, messages, max_tokens, temperature, top_p, stream: false })`
2. **Responses API**: `openai.responses.create({ input, model, temperature, top_p, previous_response_id, max_output_tokens })` with sliding window image context deletion via `openai.responses.delete(responseId)`
- Support custom headers passthrough for provider-specific extensions (e.g., `thinking: { type: 'disabled' }`)

### MCP Protocol

- Install `@modelcontextprotocol/sdk`
- **Client transports**: StdioClientTransport, SSEClientTransport, StreamableHTTPClientTransport, InMemoryTransport
- **Server**: Expose tools via stdio or HTTP using `server.json` and `smithery.yaml` configuration
- **Enhanced PATH**: For stdio servers, augment PATH with platform-specific directories (`/opt/homebrew/bin`, `~/.nvm/current/bin`, `~/.cargo/bin`, etc.)

### Browser Automation

- Install `puppeteer-core`
- Implement `LocalBrowser` (launches browser via detected executable path) and `RemoteBrowser` (connects via CDP WebSocket)
- Inject anti-detection scripts: override `navigator.webdriver`, `window.chrome`, shadow DOM `attachShadow` to `open` mode
- Inject `buildDomTree.js` script via `evaluateOnNewDocument` for DOM element extraction

### Telemetry (UTIO)

- Initialize `new UTIO(endpoint)` with POST to endpoint
- Send events: `appLaunched` (platform, osVersion, screenWidth, screenHeight), `sendInstruction` (instruction text), `shareReport` (lastScreenshot, report URL)

## 15. Core Workflows & Business Logic

### GUI Agent Execution Loop

1. User provides natural language instruction via CLI flag, Web UI input, or Electron window
2. `GUIAgent` initializes with operator (screenshot + execute) and model (invoke) configurations
3. Agent enters loop (max 25 iterations by default):
   a. Call `operator.screenshot()` to capture current screen/browser state
   b. Preprocess image: resize to `maxPixels` based on `UITarsModelVersion` (V1_0: `MAX_PIXELS_V1_0`, V1_5: `MAX_PIXELS_V1_5`, Doubao: `MAX_PIXELS_DOUBAO`)
   c. Build OpenAI messages from conversation history + compressed images via `convertToOpenAIMessages()`
   d. Call `model.invoke()` which internally calls `invokeModelProvider()` → OpenAI API
   e. Parse prediction via `actionParser({ prediction, factor, screenContext, scaleFactor, modelVer })`
   f. Emit `GUIAgentData` via `onData` callback (status, prediction, parsed actions, cost metrics)
   g. For each parsed prediction, call `operator.execute({ prediction, parsedPrediction, screenWidth, screenHeight, scaleFactor, factors })`
   h. Wait `loopIntervalInMs` then repeat
4. On error, emit via `onError` callback with `GUIAgentError`
5. On completion (agent returns `done` action or max loops), emit final status

### Browser Agent Step Execution

1. Get browser state: `page.getState()` → screenshot + DOM element tree + selector map
2. Build message with current state, previous actions memory, and user instruction
3. Invoke LLM to get `AgentOutput` containing `current_state` (page_summary, evaluation, memory, next_goal) and `action[]` array
4. Execute each action in order: click_element, input_text, scroll_down, go_to_url, switch_tab, etc.
5. Wait `waitBetweenActions` (default 1s) between actions
6. Collect `ActionResult` for each action (isDone, extractedContent, error)
7. Emit `AgentEvent` with actor, state, details, browserState

### MCP Server Activation Flow

1. Construct `MCPClient` with array of `MCPServer` configs (name, url/command, status)
2. Call `client.init()` → loads all servers with status `activate`
3. For each server: create `Client` instance, determine transport type:
   - URL-based: create StreamableHTTPClientTransport or SSEClientTransport
   - Command-based: create StdioClientTransport with enhanced PATH
   - Built-in: use InMemoryTransport with paired server/client
4. Call `client.connect(transport)` and store in `activeServers` map
5. List tools via `client.listTools()`, apply allow/block filters using `minimatch`
6. Expose filtered tools as `MCPTool[]` with generated `id` and `serverName`

### create-new-mcp Scaffolding Flow

1. Parse CLI args with `mri` (target dir, template, overwrite)
2. Prompt for project name via `@clack/prompts` if not provided
3. Handle existing directory (overwrite, cancel, or ignore)
4. Validate package name with regex `/^(?:@[a-z\d\-*~][a-z\d\-*._~]*\/)?[a-z\d\-~][a-z\d\-._~]*$/`
5. Copy `template-default/` files to target directory
6. Replace `{{variable_name}}` and `%%variable_name%%` patterns in all files
7. Output next steps: `cd <dir>`, `<pkg-manager> install`, `<pkg-manager> run dev`

## 16. Validation & Error Handling

### Zod Schemas

Use the schemas defined in Section 6 for all MCP tool inputs, agent action outputs, and IPC message validation. Every tool call argument must be validated against its corresponding Zod schema before execution.

### Error Response Format

```typescript
{ error: string; code: string; field?: string }
```

### Agent Error Handling

- Track `consecutiveFailures` in `AgentContext`; abort after `maxFailures` (default: 3)
- Truncate error messages to `maxErrorLength` (default: 400 chars) before feeding back to LLM
- Implement retry with configurable `RetryConfig` (maxRetries, onRetry callback) for model, screenshot, and execute operations
- On VLM response error (empty prediction), construct error with `name: 'vlm response error'` and `stack: JSON.stringify(result)`

### Loading States

- Visualizer: Show `<Empty>` component when no dump loaded; Dragger upload prompt for file input
- EnvConfig: Show "No config" with setup button when `Object.keys(config).length === 0`

Source: https://papers.ahmad.wtf/build-prompts/agent-tars-ui-tars-monorepo/index.mdx
