AVCodex agents can generate and edit images directly in the chat conversation. Users describe what they want and the agent creates it. No separate tools or plugins required.
AVCodex supports five image generation models across four providers. Each has different strengths, costs, and capabilities.
| Model | Provider | Best For | Cost per Image | Max Resolution |
|---|---|---|---|---|
| Gemini 3 Pro | Highest quality, text rendering, multi-reference editing | ~$0.17 | 4K | |
| GPT Image 1.5 | OpenAI | Precise instruction-following edits | ~$0.04 | 4096x4096 |
| FLUX.1 Kontext Pro | Black Forest Labs | Conversational editing, preserving the original image | ~$0.05 | 1024x1024 |
| Stability AI SD 3.5 | Stability AI | Inpainting, search-and-replace, background removal | ~$0.05 | 1024x1024 |
| Gemini 2.5 Flash | Fast, affordable, high-volume generation | ~$0.003 | Dynamic |
Note: GPT Image 1.5 is the default for new agents. Gemini 3 Pro is the recommended premium option when you need the highest quality output (rendering equipment labels in a rack diagram, for example, where text fidelity matters).
Model Capabilities#
Not every model supports every operation:
- Generation: Create images from text descriptions. All models support this.
- Editing: Modify an existing image based on instructions. Supported by all models, though Gemini 2.5 Flash has limited editing.
- Inpainting: Fill in or replace specific regions of an image. Supported by Gemini 3 Pro and Stability AI.
- Blending: Combine multiple reference images into a new composition. Supported by Gemini 3 Pro and Gemini 2.5 Flash.
Image generation is enabled by default for all agents on all tiers. To choose a model or adjust settings:
Open Build Settings.
Open your agent in the Builder and click the Build tab.
Find the Image Model Section.
Scroll to the Image Model card. You'll see a grid of available models.
Select a Model.
Click the model you want to use. The selected model is highlighted with your brand color.
Adjust Model Settings.
Each model exposes different configuration options. Adjust them below the model grid after selecting a model.
Save.
Click Save at the top of the Build tab. The new model is used for all future image-generation requests.
Each model offers different configuration parameters. Users don't see these. They're builder-level defaults applied to every generation.
GPT Image 1.5 (OpenAI)#
- Quality: Low (fastest, cheapest), Medium (balanced, default), or High (best quality). Higher quality costs more tokens.
- Size: 1024x1024 (square), 1536x1024 (landscape), or 1024x1536 (portrait).
- Background: Auto (model decides), Transparent (PNG only, useful for equipment icons or signal-flow markers), or Opaque (solid background).
Gemini 3 Pro#
- Aspect Ratio: 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, or 2:3. The model adapts composition to fit.
FLUX.1 Kontext Pro#
- Aspect Ratio: Same options as Gemini 3 Pro.
- Prompt Upsampling: When enabled, the model automatically enhances the user's prompt for better results.
Stability AI SD 3.5#
- Aspect Ratio: Similar options. Only applies to new generations. Edits preserve original dimensions.
- Negative Prompt: Describe elements to exclude (e.g. "blurry, low quality, watermark"). Steers the model away from unwanted content.
- Edit Strength: Slider from 0.1 to 1.0. Low values make subtle tweaks. High values allow dramatic transformations. Default 0.7.
Users interact with image generation through natural language. They don't need to know which model is configured. They just ask.
Generating New Images#
Users describe what they want and the agent calls the generateImage tool automatically:
- "Create a high-level signal-flow diagram for a typical 12-person huddle room with a Logitech Rally Bar and a Q-SYS Core."
- "Generate a mockup of a corporate boardroom with dual 98-inch displays and a center table mic array."
- "Draw an icon set for a control system status dashboard: green check, yellow warning, red error."
The generated image appears inline in the chat conversation.
Editing Existing Images#
Users can upload an image and ask for changes. The agent uses the uploaded image as a reference:
- "Mark up this rack elevation. Highlight the DSP and label it." (after uploading a CAD rack drawing).
- "Remove the background from this product shot of the QSC Core 110f." (after uploading a manufacturer photo).
- "Add a callout pointing to the input gain knob in this photo of the Yamaha mixer."
There's also a dedicated editImage tool that automatically resolves the most recently uploaded image, so users can say "change the color to red" without re-uploading.
Multi-Image Blending#
With Gemini models, users can upload multiple reference images and ask the agent to combine them:
- "Blend these two photos into a single composition."
- "Take the rack from image 1 and put it in the room from image 2."
Tip: Gemini 3 Pro supports up to 14 reference images in a single request, which makes it the best choice for complex multi-image workflows (combining a room photo, a CAD drawing, and a brand style sheet, for example).
If your agent doesn't need image generation, disable it in the Build tab. When disabled, the generateImage and editImage tools are removed from the agent's toolset entirely. Users won't be able to request images.
Image generation is billed per image through your organization's Stripe Token Billing balance. The cost per image depends on the model selected (see the table above). Costs shown in the Builder include a 30% platform markup over the raw provider cost.
*AVCodex · Your AV expertise. Amplified by AI.*