AICOMBO
EN
  • CN
  • EN
EN
  • CN
  • EN
EN
  • CN
  • EN
  1. Audio
  • User Guide
  • Audio
    • Text to Speech
      POST
    • Gemini Native
      POST
  • Chat
    • Native OpenAI Format
      • ChatCompletions Format
      • Responses Format
      • Responses Async Task Retrieval
      • Responses Async Task Deletion
      • Responses Async Task Retrieval (Streaming)
    • Native Gemini Format
      • Gemini Chat
    • Native Vertex AI Format
      • Gemini Chat
    • Native Anthropic Format
      • Native Claude Format
      • Native Claude Format - Token Count
  • Completions
    • Native OpenAI Format
      POST
  • Images
    • Gemini Format
      • Responses Format
      • OpenAI Compatible
  • Models
    • List Models
      • Native OpenAI Format
      • Native Gemini Format
  1. Audio

Text to Speech

POST
https://api.aicombo.com/v1/audio/speech
Convert text into audio

请求参数

Header 参数

Body 参数application/json必填

示例
{
    "model": "gemini-2.5-pro-preview-tts",
    "input": "Alan Mathison Turing was elected a Fellow of King's College, Cambridge in 1935; in 1936 he proposed a universal model of a logical machine known as the Turing machine; in 1938 he obtained his PhD from Princeton University; in 1939 he began working for the British military, during which he cracked the German Enigma cipher machine and the Tunny cipher machine, accelerating the Allied victory in World War II; in 1946 he was awarded the Order of the British Empire; from 1945 to 1948 he led research on the Automatic Computing Engine (ACE) at the National Physical Laboratory in Teddington, London; in 1948 he became a senior lecturer at the University of Manchester and assistant to the head of the Automatic Digital Machine (Madam) project; in 1949 he became deputy director of the Computing Laboratory at the University of Manchester; in 1950 he proposed the possibility that machines could think and the concept of the 'Turing test'; in 1951 he was elected a Fellow of the Royal Society; in 1954 he died after eating an apple laced with cyanide, at the age of 41.",
    "voice": "Puck"
}

请求示例代码

Shell
JavaScript
Java
Swift
Go
PHP
Python
HTTP
C
C#
Objective-C
Ruby
OCaml
Dart
R
请求示例请求示例
Shell
JavaScript
Java
Swift
cURL
curl --location 'https://api.aicombo.com/v1/audio/speech' \
--header 'Authorization: Bearer ' \
--header 'Content-Type: application/json' \
--data '{
    "model": "gemini-2.5-pro-preview-tts",
    "input": "Alan Mathison Turing was elected a Fellow of King'\''s College, Cambridge in 1935; in 1936 he proposed a universal model of a logical machine known as the Turing machine; in 1938 he obtained his PhD from Princeton University; in 1939 he began working for the British military, during which he cracked the German Enigma cipher machine and the Tunny cipher machine, accelerating the Allied victory in World War II; in 1946 he was awarded the Order of the British Empire; from 1945 to 1948 he led research on the Automatic Computing Engine (ACE) at the National Physical Laboratory in Teddington, London; in 1948 he became a senior lecturer at the University of Manchester and assistant to the head of the Automatic Digital Machine (Madam) project; in 1949 he became deputy director of the Computing Laboratory at the University of Manchester; in 1950 he proposed the possibility that machines could think and the concept of the '\''Turing test'\''; in 1951 he was elected a Fellow of the Royal Society; in 1954 he died after eating an apple laced with cyanide, at the age of 41.",
    "voice": "Puck"
}'

返回响应

🟢200Success
application/octet-stream
Audio stream. The two model families behave differently:
[OpenAI models tts-1 / tts-1-hd / gpt-4o-mini-tts]
Returns the corresponding format according to response_format; the Content-Type is passed through from upstream (the actual response header is authoritative):
mp3 → audio/mpeg
opus → audio/opus
aac → audio/aac
flac → audio/flac
wav → audio/wav
pcm → audio/pcm (24000Hz, 16-bit, mono)
When stream_format=sse, the audio is returned in chunks as an SSE event stream.
[Gemini speech models gemini-2.5-flash-preview-tts / gemini-2.5-pro-preview-tts]
Ignores response_format and always returns 24kHz / mono / 16-bit raw PCM (no file header); the Content-Type is the upstream audio/L16;rate=24000.
The client must add a WAV header itself or play it as raw PCM; it cannot be saved and played directly as mp3. SSE streaming is not supported.
上一页
User Guide
下一页
Gemini Native
Built with