Skip to content

Text-to-Speech

The Arc XP Audio API can generate spoken audio from text input by using text-to-speech (TTS). This is useful for producing narrated versions of articles, accessibility audio, or any content where you need synthesized speech.

Prerequisites

1. List available voices

This endpoint returns the full voice catalog available to your account, which is broader than the subset your organization has configured for TTS:

Terminal window
curl -H "Authorization: Bearer YOUR_API_TOKEN" \
https://api.[org].arcpublishing.com/audiocenter/api/editorial/v1/settings/voices

The response includes metadata for each voice:

{
"voices": [
{
"id": "voice_abc123",
"name": "Rachel",
"use_case": "narration",
"gender": "female",
"accent": "american",
"age": "young",
"description": "A clear, warm voice ideal for news narration.",
"preview_url": "https://...",
"supported_languages": ["en", "es", "fr"]
}
]
}

Browse this list to find an id you want to use, but note it before generating speech: only a voice already added to your organization’s settings is valid for previewing or generating speech. A handful of preset voices come configured by default; to use any other catalog voice, add it first. See Configure Voices.

2. Preview a voice

Before creating a full audio clip, you can generate a short preview (~10 seconds) to audition a voice; this preview does not create an audio clip record.

Terminal window
curl -X POST https://api.[org].arcpublishing.com/audiocenter/api/editorial/v1/tts/preview \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"input_text": "This is a short preview of what this voice sounds like.",
"voice_id": "voice_abc123"
}'

The response has a uri pointing to the generated preview audio file:

{
"uri": "https://..."
}

3. Generate a full audio clip

Once you’ve chosen a voice, create an audio clip with TTS in a single request. The request creates the clip record and starts speech synthesis:

Terminal window
curl -X POST https://api.[org].arcpublishing.com/audiocenter/api/editorial/v1/clips/tts \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"title": "Article Narration: Breaking News Story",
"description": "TTS narration of the breaking news article.",
"tags": ["narration", "news"],
"input_text": "The full text of the article you want narrated goes here. It can be up to 30,000 characters.",
"voice_id": "voice_abc123"
}'

The API returns 202 Accepted with the new clip’s ID and a Location header:

{
"content_id": "abc123def456"
}

Use that returned content_id as the {audio_id} placeholder in later clip endpoints, including publish.

Monitor progress with SSE

TTS generation is asynchronous and streams progress over server-sent events (SSE). To receive real-time updates, include the Accept: text/event-stream header:

Terminal window
curl -X POST https://api.[org].arcpublishing.com/audiocenter/api/editorial/v1/clips/tts \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Accept: text/event-stream" \
-H "Content-Type: application/json" \
-d '{
"title": "Article Narration",
"input_text": "The full text goes here...",
"voice_id": "voice_abc123"
}'

The SSE stream emits progress notifications such as tts_started, tts_generating_speech, encoding_started, and encoding_complete (or tts_failed / encoding_failed).

4. Publish the clip

Once the clip reaches ready state, publish it to make it available for delivery:

Terminal window
curl -X POST https://api.[org].arcpublishing.com/audiocenter/api/editorial/v1/clips/{audio_id}/publish \
-H "Authorization: Bearer YOUR_API_TOKEN"

Pronunciation dictionaries

If your content includes names, technical terms, or brand names that the TTS engine mispronounces, you can configure a pronunciation dictionary at the organization level.

Pronunciation rules use IPA (International Phonetic Alphabet) phonemes to define how to pronounce specific words.

Configure pronunciation rules

Terminal window
curl -X PATCH https://api.[org].arcpublishing.com/audiocenter/api/editorial/v1/settings/ \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"tts_settings": {
"pronunciation_dictionary": {
"rules": [
{
"grapheme": ["Arc XP", "ArcXP"],
"phoneme": "ɑːrk ɛks piː"
},
{
"grapheme": ["GIF"],
"phoneme": "ɡɪf"
}
]
}
}
}'

Each rule maps one or more grapheme strings (the text as written) to a phoneme (how to speak it). The dictionary applies automatically to all future TTS operations for your organization.

Configure voices

You can also configure which voices your organization can use through the settings endpoint:

Terminal window
curl -X PATCH https://api.[org].arcpublishing.com/audiocenter/api/editorial/v1/settings/ \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"tts_settings": {
"voices": [
{ "voice_id": "voice_abc123" },
{ "voice_id": "voice_def456" }
]
}
}'

TTS fields reference

Create TTS clip (POST /clips/tts)

FieldRequiredDescription
titleYesClip title (3–500 characters).
input_textYesThe text to synthesize (up to 30,000 characters).
voice_idYesID of a voice configured in your organization’s settings (see Configure Voices); not just any id from /settings/voices.
descriptionNoClip description (up to 4,000 characters).
tagsNoTags for organization and filtering.
circulationNoSites this clip belongs to.

Preview voice (POST /tts/preview)

FieldRequiredDescription
input_textYesSample text to synthesize (up to 500 characters, ~10 seconds).
voice_idYesID of the voice to audition.