Acme
AI Agent

Testing

Run the agent, tool, storage, and browser checks before shipping.

The agent ships with deterministic checks for tools, persistence, authorization, and the chat UI. Run them before you deploy changes to the agent or its tools.

Unit and procedure checks

From the repository root:

pnpm --filter @repo/agent test
pnpm --filter @repo/api test
pnpm --filter @repo/storage test

These cover tool execution, cancellation, invalidation, authorization, and persistence without calling a paid model provider.

Model-selection tests use controlled catalogs and Pi faux providers. They cover independent model/thinking controls, server allowlisting, provider credentials, per-run execution metadata, model-aware retries after catalog changes, native model/thinking transitions, and persistence failure before dispatch. Mixed-history tests exercise text-only continuation, saved image editing, original file bytes, and selected-model compaction. Do not pin the shipped model labels or catalog size.

Tool-diagnostic checks use local HTTP fixtures to verify that provider rejection bodies and request IDs reach server logs while thrown errors stay safe for the model. They also cover HTTP-200 error envelopes, bounded/truncated responses, credential and image-data redaction, concurrent tool-call attribution, and underlying application-procedure failures. See tool failure diagnostics for where to find the logs.

Diagnostic tests clear ambient credential values before installing their own fixtures, so a buyer's keys cannot accidentally redact expected fixture text. They compare log fields rather than JSON spacing or our error wording. Response-bound checks supply small test-owned byte budgets, and signed-URL redaction is exercised on a failing download.

Text-file checks cover configurable filename matching, strict UTF-8, original BOM/line-ending preservation, byte/digest integrity, admission failure before execution, retry identity, and current authorization on downloads. Session tests capture actual model context to verify file bodies reach text-only models without being copied into saved attachment records. Compacted-away references are not reintroduced; missing or oversized active files fail explicitly. Use fixture policies rather than asserting the shipped allowlist or tuning defaults.

What to test when you customize the agent

The bundled tests protect the engine so you can build on it. You should not need to rewrite them when you rename your product, edit the system prompt, replace the sample skills, change attachment limits, or select a different retrieval provider. They use their own fixtures rather than asserting your product's wording or configuration.

Unit tests use controlled credentials, bucket names, file policies, and limits. Browser and S3 file workflows select a supported filename from the running policy; they require file uploads to be enabled and enough capacity for their small fixture. Browser controls and application errors are located through the translation catalog, while deterministic model responses and user-entered fixture text remain explicit. Screenshots are diagnostic captures, not visual baselines.

When adding coverage, assert source ordering, parsed data, original bytes, authorization, and request outcomes. Avoid pinning search-tuning defaults, generated multipart filenames, exact error sentences, demo-dashboard copy, or text-size icons. Error tests must still reject failed operations and check that credentials and upstream diagnostics are not exposed. Keep exact protocol fields and byte comparisons where changing them would break the workflow.

Your changeWhat to check
Prompts or skillsTry representative questions in chat. Check answer quality, valid skill metadata, and whether the agent chooses the right tools. Passing engine tests does not measure answer quality.
An application toolTest its procedure's validation, permissions, tenant scope, and side effects. Use the existing notification procedure tests as an example.
A retrieval providerCheck its request and result mapping, then try a query against your actual knowledge base. Local HTTP fixtures do not validate your provider credentials or indexed content.
Conversation storage or runtime behaviorRun the agent suite. If you change the storage backend or persistence protocol, also run the S3 checks below.
Chat UI behaviorRun the relevant browser workflow below.

File-specific tests sit beside their implementation, such as session.ts and session.test.ts. Cross-module checks live in the owning tests/ directory. Keep the persistence, isolation, and cancellation tests: they catch lost conversations, cross-user access, and repeated actions after a failure.

Storage and database checks

Start the local services, generate the database client, and apply migrations:

pnpm install
docker compose up -d postgres seaweedfs
pnpm --filter @repo/database generate
pnpm --filter @repo/database migrate:deploy

Then run the S3 and agent storage checks:

pnpm exec dotenv -e .env.local -- pnpm --filter @repo/storage test:s3
pnpm exec dotenv -e .env.local -- pnpm --filter @repo/agent test:s3
pnpm exec dotenv -e .env.local -- pnpm --filter @repo/api test:agent-s3

The checks use a private test bucket that must reject anonymous reads and support conditional writes.

The storage contract also checks non-UTF8 binary round trips, byte limits, cancellation, and compatibility with string reads. Agent unit tests cover image decoding, validated references, immutable scoped assets, ambiguous-write reconciliation, and separation from conversation discovery. They use small image fixtures and controlled limits, with no paid image-provider calls.

The authenticated API/S3 contract includes a mixed image/text-file conversation: it checks original file bytes, denied cross-user downloads, model-visible file contents after a fresh-process restore, and subsequent image editing. The browser attachment flow checks removable chips, history restoration, original downloads, and unsupported-file recovery in both placements.

Browser checks

Run each placement separately:

AGENT_ENABLED=true AGENT_CHAT_PLACEMENT=native pnpm --filter saas e2e:agent
AGENT_ENABLED=true AGENT_CHAT_PLACEMENT=floating pnpm --filter saas e2e:agent
AGENT_ENABLED=false pnpm --filter saas e2e:agent

These tests use a deterministic model provider, real notification procedures, PostgreSQL, and private S3 storage. No model API key is required.

For the MCP Server checks, see MCP setup.

On this page