Skip to main content
POST
Gemini image models do not use /v1/images/*. They generate images through the same generateContent endpoint as text models, and return the bytes as an inline part alongside any text the model produced. For GPT image models, use POST /v1/images/generations instead. See Nano Banana for worked text-to-image and image-to-image examples.
Model IDs change over time. List the ones available to your account with GET /v1/models.

Endpoint

Authenticate with x-goog-api-key, as on every Gemini route. See Gemini overview.
array
required
Conversation turns. For text-to-image a single user turn with one text part is enough; for editing, add the source image as an inline part in the same turn.
object
Standard Gemini generation settings. Image models also accept sizing controls here; the accepted aspect ratios and resolutions differ per model, so check the model’s own card before relying on a value.
array
Generated responses.
object
Token counts. Generated images are billed as image output tokens, which dominate the cost — see Pricing.

Reading the response

Two habits save trouble:
  • Iterate over every part. Ask for an illustrated explanation and you get text and image parts interleaved. Taking only the first part silently drops content.
  • Expect a text-only response sometimes. If the prompt is refused on safety grounds the response still returns 200 with a text part and no inlineData. Check for the image part before decoding.
Images generated by Gemini carry a SynthID watermark.

Errors

Gemini routes return the Gemini error shape, not the OpenAI one. See Errors for the full mapping.
Malformed contents, an unsupported mimeType on an inline part, or a sizing value the model does not accept.
Missing or wrong x-goog-api-key.