Image generation
An agent with the generate_image tool can produce an image from a text prompt. The result is saved as a PNG artifact and shows up in chat like any other attachment.
Setup
Two things are required.
- An AI-image provider. Image generation runs against the OpenAI Images API contract, served by an
openaiprovider or an OpenAI-compatible provider with a base URL (that is, a self-hosted backend). Nakama uses your default provider when it matches, otherwise the first eligible provider configured. The key comes from that provider, or fromOPENAI_API_KEY. - A model, set in AI Providers. Pick Image generation model in the same card group as the vision and transcription models. The choices are
gpt-image-2against an OpenAI provider, or any model of an OpenAI-compatible provider, stored as<providerId>::<modelId>. See Self-hosted models.
Until a model is set, the tool refuses with Configure an image generation model in Settings before generating images. The card shows Not configured in that state, and No image provider when the provider is the missing half.
Assign the generate_image tool to a profile the same way as any other builtin. See Builtin tools.
Using it
Ask in plain language. The agent calls generate_image with a prompt and, optionally, a filename and a size.
To be explicit, type @ in the web chat composer, pick @image, and describe the picture. A message tagged @image tells the agent to call generate_image straight away. When the profile does not have the tool, the agent says so and points to the tool assignment and the Settings model instead of answering in text. The tag only counts at the start of a word, so an address such as me@image.dev stays plain text.
| Parameter | Required | Notes |
|---|---|---|
prompt | yes | What to draw |
filename | no | Output name under artifacts/, defaults to a generated .png |
size | no | 1024x1024 (default), 1024x1536, 1536x1024, or auto |
The model is deliberately not a parameter. It comes from AI Providers, so a prompt cannot talk the agent into a model the workspace has not allowlisted. An unsupported size is rejected with the allowed list rather than silently rounded to the nearest one.
Where the image goes
Each generated image is written under the profile workspace:
~/.nakama/orgs/{orgId}/profiles/{profileId}/artifacts/The image appears in Browse Files → Artifacts with its type, size, and timestamp detected automatically. When the call happens inside a chat session, the image is also saved as a session attachment, which makes it render inline in web chat. Web chat and channel bridges recognize generate_image output, so Telegram and Discord turns also get the picture. The tool refuses generated files over 5 MB.
Channel behavior follows each channel's artifact rules. Telegram receives a share link after the save and caps a native document attachment at 5 MB. Discord sends supported native attachments up to 8 MB and otherwise falls back to the share link. WhatsApp automatically sends the newly saved PNG as a native document up to 16 MB. See Artifacts on channels for the delivery details.
See Artifacts and files for previewing, editing, and publishing what comes out.
Cost
Every generation records usage through the same token pricing bridge as chat, so image spend shows up next to model spend rather than being invisible. When the API returns a usage breakdown, that is what gets recorded. When it does not, Nakama estimates: input tokens from the prompt length, output tokens from the requested size. Estimated figures are recorded as estimates, not passed off as measured.
Limits
gpt-image-2on an OpenAI provider, or a model of an OpenAI-compatible provider with a base URL that speaks the OpenAI Images API (see Self-hosted models). There is no Stability or Replicate path.- No image editing or variations. Generation from a prompt only.
- No streaming or progress. A generation is one tool call that returns when the image is ready. A provider that has not answered completely within 180 seconds fails with a timeout error, and stopping the chat aborts the request.
- One PNG per call, up to 5 MB, with
1024x1024,1024x1536,1536x1024, orautooutput size.