VoiceDeveloperWhy Voice‑First Matters in Modern Web Apps If you’ve ever built a dashboard or a chat...
If you’ve ever built a dashboard or a chat interface, you know how quickly visual clutter can overwhelm users. Voice‑first experiences cut through that noise, letting people listen instead of read and talk instead of type. Whether you’re creating an accessibility‑focused app, a hands‑free productivity tool, or just adding a fun “read‑me‑the‑news” button, a solid text‑to‑speech (TTS) engine is the backbone of any voice‑enabled UI.
Enter ElevenLabs – a TTS service that delivers natural‑sounding, expressive speech and even supports voice cloning. In this walkthrough we’ll wire up ElevenLabs to a React front‑end, giving you a reusable component that can read any string of text aloud with just a click.
Pro tip: If you’re looking for an API that feels like a modern SaaS product (clear docs, generous free tier, and low‑latency streaming), ElevenLabs is a great default choice. Grab your API key here: https://try.elevenlabs.io/kr07zfuqn1bp
.env file:
# .env
REACT_APP_ELEVENLABS_API_KEY=your_secret_key_here
Security note: Never commit your
.envfile. Use tools likedotenvor Vite’s built‑in env handling to inject the variable at build time.
ElevenLabs exposes a simple POST /v1/text-to-speech/{voice_id} endpoint. The request body is JSON and the response is an audio stream (MP3 by default). Here’s the raw curl version, just to illustrate the shape of the request:
curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/EXAMPLE_VOICE_ID" \
-H "xi-api-key: $ELEVENLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"text": "Hello, this is a demo of ElevenLabs in a React app!",
"model_id": "eleven_monolingual_v1",
"voice_settings": {
"stability": 0.75,
"similarity_boost": 0.85
}
}' \
--output output.mp3
The same call can be made from the browser using fetch. Because browsers block cross‑origin binary streams by default, we’ll pipe the response into a Blob and hand it to an HTMLAudioElement.
Let’s encapsulate the fetch logic in a custom hook called useElevenLabsTTS. It will:
play function that triggers the request and plays the audio.
// src/hooks/useElevenLabsTTS.ts
import { useState, useCallback } from "react";
type TTSOptions = {
voiceId?: string; // defaults to a premium voice if omitted
stability?: number;
similarityBoost?: number;
};
export function useElevenLabsTTS(
text: string,
{ voiceId = "EXAMPLE_VOICE_ID", stability = 0.75, similarityBoost = 0.85 }: TTSOptions = {}
) {
const [loading, setLoading] = useState(false);
const [error, setError] = useState<string | null>(null);
const play = useCallback(async () => {
setLoading(true);
setError(null);
try {
const resp = await fetch(
`https://api.elevenlabs.io/v1/text-to-speech/${voiceId}`,
{
method: "POST",
headers: {
"Content-Type": "application/json",
"xi-api-key": import.meta.env.VITE_ELEVENLABS_API_KEY,
},
body: JSON.stringify({
text,
model_id: "eleven_monolingual_v1",
voice_settings: { stability, similarity_boost: similarityBoost },
}),
}
);
if (!resp.ok) {
const err = await resp.text();
throw new Error(`ElevenLabs error: ${err}`);
}
const audioBlob = await resp.blob();
const audioUrl = URL.createObjectURL(audioBlob);
const audio = new Audio(audioUrl);
audio.play();
} catch (e: any) {
setError(e.message);
} finally {
setLoading(false);
}
}, [text, voiceId, stability, similarityBoost]);
return { play, loading, error };
}
Note: The hook uses Vite’s
import.meta.envconvention. If you’re on Create‑React‑App, replace it withprocess.env.REACT_APP_ELEVENLABS_API_KEY.
Now we can build a tiny component that lets users type a sentence and hear it back instantly.
// src/components/VoiceReader.tsx
import { useState } from "react";
import { useElevenLabsTTS } from "../hooks/useElevenLabsTTS";
export function VoiceReader() {
const [input, setInput] = useState("Hello, world! This is ElevenLabs speaking.");
const { play, loading, error } = useElevenLabsTTS(input);
return (
<div style={{ maxWidth: 600, margin: "2rem auto", padding: "1rem", border: "1px solid #eaeaea", borderRadius: 8 }}>
<h2>🗣️ Voice‑Enabled UI Demo</h2>
<textarea
rows={3}
style={{ width: "100%", fontSize: "1rem", marginBottom: "0.5rem" }}
value={input}
onChange={(e) => setInput(e.target.value)}
/>
<button
onClick={play}
disabled={loading}
style={{
padding: "0.5rem 1rem",
background: loading ? "#999" : "#0070f3",
color: "#fff",
border: "none",
borderRadius: 4,
cursor: loading ? "default" : "pointer",
}}
>
{loading ? "Generating…" : "Speak"}
</button>
{error && <p style={{ color: "red", marginTop: "0.5rem" }}>{error}</p>}
</div>
);
}
Drop <VoiceReader /> anywhere in your app (e.g., inside App.tsx). When the user clicks Speak, the hook contacts ElevenLabs, streams back an MP3, and the browser plays it automatically.
One of ElevenLabs’ standout features is voice cloning. If you have a short sample (≈30 seconds) of a speaker, you can upload it via the API and receive a new voice_id that mimics that voice. The workflow is:
/v1/voices/add.
voice_id in the TTS request.Here’s a quick Node script to create a clone:
// cloneVoice.js (run with node)
const fetch = require("node-fetch");
require("dotenv").config();
async function createClone(samplePath) {
const sample = require("fs").createReadStream(samplePath);
const resp = await fetch("https://api.elevenlabs.io/v1/voices/add", {
method: "POST",
headers: {
"xi-api-key": process.env.ELEVENLABS_API_KEY,
},
body: sample,
});
const { voice_id, status_url } = await resp.json();
console.log(`Clone created! Voice ID: ${voice_id}`);
console.log(`Check status at: ${status_url}`);
}
createClone("./my-sample.wav");
Once you have the new voice_id, simply pass it to the useElevenLabsTTS hook:
const { play } = useElevenLabsTTS(input, { voiceId: "YOUR_CLONE_ID" });
Caution: Voice cloning should respect privacy and consent. Only clone voices you own or have explicit permission to use.
| Issue | Fix |
|---|---|
| Delay before audio plays | Use the stream endpoint (/v1/text-to-speech/{voice_id}/stream) to start playback as soon as the first chunk arrives. |
| Long sentences cause timeouts | Break long text into 200‑character chunks and queue them. |
| Multiple concurrent requests | Keep a single Audio instance and cancel any ongoing playback before starting a new one (audio.pause(); audio.currentTime = 0;). |
| Accessibility | Pair the spoken output with aria-live regions so screen readers announce the same content. |
When you move from local dev to production, remember to:
* for production keys, but a proxy gives you more control.A minimal Vercel function could look like:
// api/tts.js
export default async function handler(req, res) {
const { text, voiceId } = req.body;
const apiResp = await fetch(
`https://api.elevenlabs.io/v1/text-to-speech/${voiceId}`,
{
method: "POST",
headers: {
"Content-Type": "application/json",
"xi-api-key": process.env.ELEVENLABS_API_KEY,
},
body: JSON.stringify({
text,
model_id: "eleven_monolingual_v1",
}),
}
);
const audioBlob = await apiResp.blob();
res.setHeader("Content-Type", "audio/mpeg");
res.send(Buffer.from(await audioBlob.arrayBuffer()));
}
Your front‑end would now hit /api/tts instead of the raw ElevenLabs endpoint, keeping the key secret.
Integrating voice AI into a React app is surprisingly straightforward when you have a reliable TTS backend. With just a few lines of code you can:
All of this is powered by ElevenLabs, a service that balances quality, flexibility, and developer ergonomics. If you’ve been curious about adding a voice layer to your next project, now’s the perfect time.
Ready to give your UI a voice? Grab your free ElevenLabs API key, spin up the demo component, and start experimenting. Happy coding! 🚀
Try ElevenLabs today: https://try.elevenlabs.io/kr07zfuqn1bp