Gemini Gerando Conteúdo

POST

v1beta

models

{model}

{operator}

from google import genai

client = genai.Client(
    api_key="<COMETAPI_KEY>",
    http_options={"api_version": "v1beta", "base_url": "https://api.cometapi.com"},
)

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Explain how AI works in a few words",
)

print(response.text)

{
  "candidates": [
    {
      "content": {
        "role": "<string>",
        "parts": [
          {
            "text": "<string>",
            "functionCall": {
              "name": "<string>",
              "args": {}
            },
            "inlineData": {
              "mimeType": "<string>",
              "data": "<string>"
            },
            "thought": true
          }
        ]
      },
      "finishReason": "STOP",
      "safetyRatings": [
        {
          "category": "<string>",
          "probability": "<string>",
          "blocked": true
        }
      ],
      "citationMetadata": {
        "citationSources": [
          {
            "startIndex": 123,
            "endIndex": 123,
            "uri": "<string>",
            "license": "<string>"
          }
        ]
      },
      "tokenCount": 123,
      "avgLogprobs": 123,
      "groundingMetadata": {
        "groundingChunks": [
          {
            "web": {
              "uri": "<string>",
              "title": "<string>"
            }
          }
        ],
        "groundingSupports": [
          {
            "groundingChunkIndices": [
              123
            ],
            "confidenceScores": [
              123
            ],
            "segment": {
              "startIndex": 123,
              "endIndex": 123,
              "text": "<string>"
            }
          }
        ],
        "webSearchQueries": [
          "<string>"
        ]
      },
      "index": 123
    }
  ],
  "promptFeedback": {
    "blockReason": "SAFETY",
    "safetyRatings": [
      {
        "category": "<string>",
        "probability": "<string>",
        "blocked": true
      }
    ]
  },
  "usageMetadata": {
    "promptTokenCount": 123,
    "candidatesTokenCount": 123,
    "totalTokenCount": 123,
    "trafficType": "<string>",
    "thoughtsTokenCount": 123,
    "promptTokensDetails": [
      {
        "modality": "<string>",
        "tokenCount": 123
      }
    ],
    "candidatesTokensDetails": [
      {
        "modality": "<string>",
        "tokenCount": 123
      }
    ]
  },
  "modelVersion": "<string>",
  "createTime": "<string>",
  "responseId": "<string>"
}

Visão geral

A CometAPI oferece suporte ao formato nativo da API Gemini, dando a você acesso completo a recursos específicos do Gemini, como controle de thinking, grounding com Google Search, modalidades nativas de geração de imagem e muito mais. Use este endpoint quando precisar de capacidades que não estão disponíveis por meio do endpoint de chat compatível com OpenAI.

Início rápido

Substitua a URL base e a chave de API em qualquer SDK do Gemini ou cliente HTTP:

Configuração	Padrão do Google	CometAPI
URL base	`generativelanguage.googleapis.com`	`api.cometapi.com`
Chave de API	`$GEMINI_API_KEY`	`$COMETAPI_KEY`

Tanto os cabeçalhos x-goog-api-key quanto Authorization: Bearer são compatíveis para autenticação.

Thinking (Reasoning)

Os modelos Gemini podem realizar reasoning interno antes de gerar uma resposta. O método de controle depende da geração do modelo.

Gemini 3 (thinkingLevel)
Gemini 2.5 (thinkingBudget)

Os modelos Gemini 3 usam thinkingLevel para controlar a profundidade do reasoning. Níveis disponíveis: MINIMAL, LOW, MEDIUM, HIGH.

curl "https://api.cometapi.com/v1beta/models/gemini-3.1-pro-preview:generateContent" \
  -H "Content-Type: application/json" \
  -H "x-goog-api-key: $COMETAPI_KEY" \
  -d '{
    "contents": [{"parts": [{"text": "Explain quantum physics simply."}]}],
    "generationConfig": {
      "thinkingConfig": {"thinkingLevel": "LOW"}
    }
  }'

Os modelos Gemini 2.5 usam thinkingBudget para controle refinado em nível de token:

0 — desabilita o thinking
-1 — dinâmico (o modelo decide, padrão)
> 0 — orçamento específico de token (por exemplo, 1024, 2048)

curl "https://api.cometapi.com/v1beta/models/gemini-2.5-flash:generateContent" \
  -H "Content-Type: application/json" \
  -H "x-goog-api-key: $COMETAPI_KEY" \
  -d '{
    "contents": [{"parts": [{"text": "Solve this logic puzzle step by step."}]}],
    "generationConfig": {
      "thinkingConfig": {"thinkingBudget": 2048}
    }
  }'

Usar thinkingLevel com modelos Gemini 2.5 (ou thinkingBudget com modelos Gemini 3) pode causar erros. Use o parâmetro correto para a versão do seu modelo.

Streaming

Use streamGenerateContent?alt=sse como operador para receber Server-Sent Events à medida que o modelo gera conteúdo. Cada evento SSE contém uma linha data: com um objeto JSON GenerateContentResponse.

curl "https://api.cometapi.com/v1beta/models/gemini-2.5-flash:streamGenerateContent?alt=sse" \
  -H "Content-Type: application/json" \
  -H "x-goog-api-key: $COMETAPI_KEY" \
  --no-buffer \
  -d '{
    "contents": [{"parts": [{"text": "Write a short poem about the stars"}]}]
  }'

Instruções do sistema

Oriente o comportamento do modelo ao longo de toda a conversa com systemInstruction:

curl "https://api.cometapi.com/v1beta/models/gemini-2.5-flash:generateContent" \
  -H "Content-Type: application/json" \
  -H "x-goog-api-key: $COMETAPI_KEY" \
  -d '{
    "contents": [{"parts": [{"text": "What is 2+2?"}]}],
    "systemInstruction": {
      "parts": [{"text": "You are a math tutor. Always show your work."}]
    }
  }'

Modo JSON

Force saída JSON estruturada com responseMimeType. Opcionalmente, forneça um responseSchema para validação estrita de schema:

curl "https://api.cometapi.com/v1beta/models/gemini-2.5-flash:generateContent" \
  -H "Content-Type: application/json" \
  -H "x-goog-api-key: $COMETAPI_KEY" \
  -d '{
    "contents": [{"parts": [{"text": "List 3 planets with their distances from the sun"}]}],
    "generationConfig": {
      "responseMimeType": "application/json"
    }
  }'

Grounding com Google Search

Ative a busca na web em tempo real adicionando uma ferramenta googleSearch:

curl "https://api.cometapi.com/v1beta/models/gemini-2.5-flash:generateContent" \
  -H "Content-Type: application/json" \
  -H "x-goog-api-key: $COMETAPI_KEY" \
  -d '{
    "contents": [{"parts": [{"text": "Who won the euro 2024?"}]}],
    "tools": [{"google_search": {}}]
  }'

A resposta inclui groundingMetadata com URLs de origem e pontuações de confiança.

Exemplo de Resposta

Uma resposta típica do endpoint Gemini da CometAPI:

{
  "candidates": [
    {
      "content": {
        "role": "model",
        "parts": [{"text": "Hello"}]
      },
      "finishReason": "STOP",
      "avgLogprobs": -0.0023
    }
  ],
  "usageMetadata": {
    "promptTokenCount": 5,
    "candidatesTokenCount": 1,
    "totalTokenCount": 30,
    "trafficType": "ON_DEMAND",
    "thoughtsTokenCount": 24,
    "promptTokensDetails": [{"modality": "TEXT", "tokenCount": 5}],
    "candidatesTokensDetails": [{"modality": "TEXT", "tokenCount": 1}]
  },
  "modelVersion": "gemini-2.5-flash",
  "createTime": "2026-03-25T04:21:43.756483Z",
  "responseId": "CeynaY3LDtvG4_UP0qaCuQY"
}

O campo thoughtsTokenCount em usageMetadata mostra quantos tokens o modelo gastou em raciocínio interno, mesmo quando a saída de thinking não está incluída na resposta.

Principais Diferenças em Relação ao Endpoint Compatível com OpenAI

Recurso	Gemini nativo (`/v1beta/models/...`)	Compatível com OpenAI (`/v1/chat/completions`)
Controle de thinking	`thinkingConfig` com `thinkingLevel` / `thinkingBudget`	Não disponível
Grounding com Google Search	`tools: [\{"google_search": \{\}\}]`	Não disponível
Grounding com Google Maps	`tools: [\{"googleMaps": \{\}\}]`	Não disponível
Modalidade de geração de imagem	`responseModalities: ["IMAGE"]`	Não disponível
Cabeçalho de autenticação	`x-goog-api-key` ou `Bearer`	Apenas `Bearer`
Formato da resposta	Gemini nativo (`candidates`, `parts`)	Formato OpenAI (`choices`, `message`)

Autorizações

x-goog-api-key

string

header

obrigatório

Your CometAPI key passed via the x-goog-api-key header. Bearer token authentication (Authorization: Bearer <key>) is also supported.

Parâmetros de caminho

model

string

obrigatório

The Gemini model ID to use. See the Models page for current Gemini model IDs.

Exemplo:

"gemini-2.5-flash"

operator

enum<string>

obrigatório

The operation to perform. Use generateContent for synchronous responses, or streamGenerateContent?alt=sse for Server-Sent Events streaming.

Opções disponíveis:

generateContent,

streamGenerateContent?alt=sse

Exemplo:

"generateContent"

Corpo

application/json

contents

object[]

obrigatório

The conversation history and current input. For single-turn queries, provide a single item. For multi-turn conversations, include all previous turns.

Show child attributes

systemInstruction

object

System instructions that guide the model's behavior across the entire conversation. Text only.

Show child attributes

tools

object[]

Tools the model may use to generate responses. Supports function declarations, Google Search, Google Maps, and code execution.

Show child attributes

toolConfig

object

Configuration for tool usage, such as function calling mode.

Show child attributes

safetySettings

object[]

Safety filter settings. Override default thresholds for specific harm categories.

Show child attributes

generationConfig

object

Configuration for model generation behavior including temperature, output length, and response format.

Show child attributes

cachedContent

string

The name of cached content to use as context. Format: cachedContents/{id}. See the Gemini context caching documentation for details.

Resposta

200 - application/json

Successful response. For streaming requests, the response is a stream of SSE events, each containing a GenerateContentResponse JSON object prefixed with data:.

candidates

object[]

The generated response candidates.

Show child attributes

promptFeedback

object

Feedback on the prompt, including safety blocking information.

Show child attributes

usageMetadata

object

Token usage statistics for the request.

Show child attributes

modelVersion

string

The model version that generated this response.

createTime

string

The timestamp when this response was created (ISO 8601 format).

responseId

string

Unique identifier for this response.

Anthropic Messages

Embeddings

from google import genai

client = genai.Client(
    api_key="<COMETAPI_KEY>",
    http_options={"api_version": "v1beta", "base_url": "https://api.cometapi.com"},
)

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Explain how AI works in a few words",
)

print(response.text)

{
  "candidates": [
    {
      "content": {
        "role": "<string>",
        "parts": [
          {
            "text": "<string>",
            "functionCall": {
              "name": "<string>",
              "args": {}
            },
            "inlineData": {
              "mimeType": "<string>",
              "data": "<string>"
            },
            "thought": true
          }
        ]
      },
      "finishReason": "STOP",
      "safetyRatings": [
        {
          "category": "<string>",
          "probability": "<string>",
          "blocked": true
        }
      ],
      "citationMetadata": {
        "citationSources": [
          {
            "startIndex": 123,
            "endIndex": 123,
            "uri": "<string>",
            "license": "<string>"
          }
        ]
      },
      "tokenCount": 123,
      "avgLogprobs": 123,
      "groundingMetadata": {
        "groundingChunks": [
          {
            "web": {
              "uri": "<string>",
              "title": "<string>"
            }
          }
        ],
        "groundingSupports": [
          {
            "groundingChunkIndices": [
              123
            ],
            "confidenceScores": [
              123
            ],
            "segment": {
              "startIndex": 123,
              "endIndex": 123,
              "text": "<string>"
            }
          }
        ],
        "webSearchQueries": [
          "<string>"
        ]
      },
      "index": 123
    }
  ],
  "promptFeedback": {
    "blockReason": "SAFETY",
    "safetyRatings": [
      {
        "category": "<string>",
        "probability": "<string>",
        "blocked": true
      }
    ]
  },
  "usageMetadata": {
    "promptTokenCount": 123,
    "candidatesTokenCount": 123,
    "totalTokenCount": 123,
    "trafficType": "<string>",
    "thoughtsTokenCount": 123,
    "promptTokensDetails": [
      {
        "modality": "<string>",
        "tokenCount": 123
      }
    ],
    "candidatesTokensDetails": [
      {
        "modality": "<string>",
        "tokenCount": 123
      }
    ]
  },
  "modelVersion": "<string>",
  "createTime": "<string>",
  "responseId": "<string>"
}

Visão Geral

Referência da API

Guias de Integração

Erros

Preços e Faturamento

Suporte

Visão geral

Início rápido

Thinking (Reasoning)

Streaming

Instruções do sistema

Modo JSON

Grounding com Google Search

Exemplo de Resposta

Principais Diferenças em Relação ao Endpoint Compatível com OpenAI

Autorizações

Parâmetros de caminho

Corpo

Resposta

Visão Geral

Referência da API

Guias de Integração

Erros

Preços e Faturamento

Suporte

​Visão geral

​Início rápido

​Thinking (Reasoning)

​Streaming

​Instruções do sistema

​Modo JSON

​Grounding com Google Search

​Exemplo de Resposta

​Principais Diferenças em Relação ao Endpoint Compatível com OpenAI

Autorizações

Parâmetros de caminho

Corpo

Resposta

Visão geral

Início rápido

Thinking (Reasoning)

Streaming

Instruções do sistema

Modo JSON

Grounding com Google Search

Exemplo de Resposta

Principais Diferenças em Relação ao Endpoint Compatível com OpenAI