Image generation and editing
Configure native image tools with private conversation assets.
The embedded agent exposes two optional native tools:
image_gen: ordered text parts and optional conversation image references, plus an optional aspect ratio.image_edit: ordered text and image references. The first image is the edit target; later images are references. Edits create a new asset and preserve every original. PNG, JPEG, and WebP sources are supported; GIF attachments remain readable but cannot be used as references.
Both tools return one private saved image reference. The main conversational model interprets requests and chooses the tool; image providers do not replace it. Explicit generation and editing requests work in ordinary chat. Image intent provides additional run-local guidance from packages/agent/prompts/image-mode.md.
Both tools require at least one nonblank text part. The agent chooses relevant instructions, history excerpts, and image IDs rather than automatically forwarding the whole conversation:
{
"content": [
{ "type": "text", "text": "Edit this scene:" },
{ "type": "image", "imageId": "scene-id" },
{ "type": "text", "text": "Add the person from this photo:" },
{ "type": "image", "imageId": "person-id" },
{ "type": "text", "text": "Keep the warm evening lighting." }
]
}Gemini receives native interleaved text/image parts. The other adapters serialize all text and numbered image-position markers into one prompt, with images in matching order. xAI markers also include its native zero-based <IMAGE_0> notation. There is no extra model call to summarize context.
Use images in chat
Click the Image generation button immediately after the text-size control to turn on image mode. It expands to show an Aspect ratio dropdown, initially square (1:1), with the selected ratio displayed beside the image icon. Click the image button again to turn image mode off; the icon stays available and the ratio dropdown collapses. The plus menu contains photo attachments and tool information.
The ratio choices come from the running API's selected image tool configuration. The selected ratio is saved with each message admission and enforced for generation in that run, even if the conversational model omits it or supplies another supported ratio. Editing keeps the provider's editing behavior. Turning image mode off lets ordinary chat requests choose their own ratio again. Mode and ratio choices survive sends, conversation switches, and navigation within the same personal or organization scope during a page load. Refresh returns to normal chat. Changing controls does not stop an active run or clear your draft.
After changing the backend provider or its configured ratios, restart/redeploy the API and refresh the page to load the new options. The UI caches capabilities and does not poll; refreshing reloads them. IMAGE_MODEL alone does not discover or change supported ratios. A stale, unsupported selection is rejected before admission or paid generation.
Generated images appear in the assistant turn that created them. While generation or editing runs, both chat placements show an image-sized, softly pulsing placeholder instead of a tool-call row or tool-call count. The placeholder names the current image operation, respects reduced-motion preferences, and stops pulsing on failure or interruption. Completed images replace the pending presentation; image loading also keeps an image-shaped placeholder until the private asset is available.
In image mode, a successful saved generation or edit finishes the run after the native tool results have been persisted, without another conversational-model request. Failed tools can still be explained by the assistant. In ordinary chat, the assistant can continue with a text response; image results explain that the UI already displays the asset and that image IDs are not Markdown URLs. Download image performs a fresh authorized request. Attach a photo and ask for an edit, or ask to change an earlier generated image; saved images remain available after reopening the conversation. If loading fails, Retry image retries the read without generating another image.
The text-size trigger uses T with a subscript S, D, or L for Small, Default, or Large. Full size names remain in its menu, and sizing changes only the conversation display.
Enable and select a provider
Set enabled: true in packages/agent/tools/image/config.ts. Select one direct adapter through the static export in packages/agent/tools/image/provider/index.ts. The integration is disabled by default, with OpenAI selected. Tool discovery makes no provider requests, and credentials are read only when a tool executes.
Each provider has an editable provider/<name>/config.ts containing its default model, allowed aspect ratios, size mappings where needed, and input limits. Edit these together without changing its request implementation. For size-based providers, advertised ratios are derived from the size-table keys.
Set the server-only IMAGE_MODEL environment variable to override the selected provider's model for both generation and editing. Whitespace is trimmed; unset, empty, or whitespace-only values use that provider's configured default:
process.env.IMAGE_MODEL?.trim() || providerConfig.model;The value is a model identifier for the selected provider, not a provider name or URL. Switching providers retains the same override, so update it to an identifier that the new provider supports. An inaccessible or unsupported model fails without retrying with the default. This setting does not change the conversational model catalog or infer image capabilities: keep the provider config's ratios and limits compatible with your chosen model.
OpenAI
Set OPENAI_API_KEY. The adapter defaults to gpt-image-2.5-sunburst. Text-only generation uses the direct Images generations endpoint; reference-guided generation and editing use ordered multipart image[] inputs on edits. Settings live in provider/openai/config.ts. Your organization may need verification and model access.
Google Gemini
Select ./google and set GOOGLE_GENERATIVE_AI_API_KEY. The adapter defaults to gemini-3.1-flash-image through Interactions, with store: false and ordered explicit source bytes on each request. It does not retain a Google conversation ID. Settings live in provider/google/config.ts.
These credentials belong to the image adapter. Conversational models are selected
in packages/agent/models.ts; their picker does not change this adapter, its model,
or image ratios. A text-only conversational model can still invoke image tools on
saved references, but new chat image attachments require a vision-capable model.
Black Forest Labs
Select ./bfl and set BFL_API_KEY. The adapter defaults to flux-2-pro with references in input_image, input_image_2, and subsequent numbered fields. It submits once, polls the returned job URL, and downloads the result immediately because the URL expires after ten minutes. The deadline includes polling and download. Model, size, limit, and polling settings live in provider/bfl/config.ts.
xAI
Select ./xai and set XAI_API_KEY. The adapter defaults to grok-imagine-image-2.0. Reference-guided requests use the edits endpoint with an ordered JSON images collection of data URIs. It requests base64 output and also handles temporary output URLs through the bounded private-address-blocking downloader. Settings live in provider/xai/config.ts.
Ideogram
Select ./ideogram and set IDEOGRAM_API_KEY. The adapter defaults to ideogram-4-5. Generation accepts multipart images; precise edits send a primary image and later reference_images. Only safe results are downloaded. Generation uses configured output sizes; precise edits retain source framing. Sources outside 1:6–6:1 are rejected before submission. Settings live in provider/ideogram/config.ts.
Provider-specific input limits
Each provider config exports static limits, re-exported by its adapter. Tool discovery advertises them without network calls. Unsupported inputs fail before paid dispatch; aspect ratios are never silently substituted or cropped.
| Adapter | Input images (generation / edit) | Total text | Generation ratios |
|---|---|---|---|
| OpenAI | 16 / 16 | 4,000 characters | 1:1, 3:2, 2:3 |
| 14 / 14 | 4,000 characters | 1:1, 3:2, 2:3, 4:3, 3:4, 16:9, 9:16 | |
| BFL | 8 / 8 | 4,000 characters | 1:1, 3:2, 2:3 |
| xAI | 5 / 5 | 4,000 characters | 1:1, 3:2, 2:3, 4:3, 3:4, 16:9, 9:16 |
| Ideogram | 5 / 5 | 4,000 characters | 1:1, 3:2, 2:3 |
The adapters expose a bounded set of their models' controls. When changing an adapter's model, update its supported mappings and limits together.
Limits and recovery
The integration defaults to 64 content parts, 32 MiB combined source bytes, a 180-second overall deadline, 16 MiB output bytes, and a 16-megapixel decoded-image limit. Provider configs currently bound each source to 16 MiB / 16 megapixels and the final encoded request to 48 MiB; these are application ceilings, not claims about vendor maxima. Use the lower of shared and provider input budgets. The binary storage ceiling remains 16 MiB.
Prompt framing counts toward the text limit; repeated image occurrences count toward image and byte limits even when their private asset is resolved once. JSON budgets include Base64 expansion; multipart budgets include conservative framing overhead. All images must resolve and validate before submission. Oversized context is rejected without truncation or splitting into extra paid requests. The overall deadline also covers source resolution. Edits follow provider framing; exact untouched pixels and identical dimensions are not guaranteed across providers.
Existing saved images and historical single-image tool calls remain readable, including outputs from the removed Adobe adapter. History is neither rewritten nor replayed. New tool calls must use content; obsolete top-level prompt / imageId inputs are rejected.
Private assets use the existing agent bucket and require conditional writes. See agent setup for persistence and retention details. Each tool invocation rechecks the originating login and organization membership. Tool arguments cannot select provider URLs, credentials, storage paths, or model overrides.
The runtime does not automatically retry paid generation/edit submissions or fall back to another provider. A timeout or lost response can leave the upstream outcome unknown. Stop ends local work and polling; it cannot guarantee that the provider stops billing. A storage failure prevents a saved-image acknowledgement and stops subsequent work in that run.
Test your integration
pnpm --filter @repo/agent test
pnpm exec dotenv -e .env.local -- pnpm --filter @repo/storage test:s3The deterministic tests use controlled configuration and local HTTP fixtures. They validate request mapping, parsing, bounds, authorization, and persistence without paid API calls. They do not establish model availability for your account. With working provider credentials, separately verify one generation and a follow-up edit; report missing credentials as skipped live verification.