ElevenLabs + React: Building a Voice-Enabled UI

# webdev# javascript# tutorial# ai
ElevenLabs + React: Building a Voice-Enabled UIVoiceDeveloper

Why Voice‑First Matters in Modern Web Apps If you’ve ever built a dashboard or a chat...

Why Voice‑First Matters in Modern Web Apps

If you’ve ever built a dashboard or a chat interface, you know how quickly visual clutter can overwhelm users. Voice‑first experiences cut through that noise, letting people listen instead of read and talk instead of type. Whether you’re creating an accessibility‑focused app, a hands‑free productivity tool, or just adding a fun “read‑me‑the‑news” button, a solid text‑to‑speech (TTS) engine is the backbone of any voice‑enabled UI.

Enter ElevenLabs – a TTS service that delivers natural‑sounding, expressive speech and even supports voice cloning. In this walkthrough we’ll wire up ElevenLabs to a React front‑end, giving you a reusable component that can read any string of text aloud with just a click.

Pro tip: If you’re looking for an API that feels like a modern SaaS product (clear docs, generous free tier, and low‑latency streaming), ElevenLabs is a great default choice. Grab your API key here: https://try.elevenlabs.io/kr07zfuqn1bp


1. Getting Started – Grab an API Key

  1. Sign up at the link above.
  2. Navigate to Dashboard → API Keys and create a new key.
  3. Store that key somewhere safe – for local development you can drop it into a .env file:
# .env
REACT_APP_ELEVENLABS_API_KEY=your_secret_key_here
Enter fullscreen mode Exit fullscreen mode

Security note: Never commit your .env file. Use tools like dotenv or Vite’s built‑in env handling to inject the variable at build time.


2. The Core API Call

ElevenLabs exposes a simple POST /v1/text-to-speech/{voice_id} endpoint. The request body is JSON and the response is an audio stream (MP3 by default). Here’s the raw curl version, just to illustrate the shape of the request:

curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/EXAMPLE_VOICE_ID" \
  -H "xi-api-key: $ELEVENLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "text": "Hello, this is a demo of ElevenLabs in a React app!",
        "model_id": "eleven_monolingual_v1",
        "voice_settings": {
          "stability": 0.75,
          "similarity_boost": 0.85
        }
      }' \
  --output output.mp3
Enter fullscreen mode Exit fullscreen mode

The same call can be made from the browser using fetch. Because browsers block cross‑origin binary streams by default, we’ll pipe the response into a Blob and hand it to an HTMLAudioElement.


3. Building a Reusable React Hook

Let’s encapsulate the fetch logic in a custom hook called useElevenLabsTTS. It will:

  • Accept a piece of text and an optional voice ID.
  • Return a play function that triggers the request and plays the audio.
  • Manage loading/error states for UI feedback.
// src/hooks/useElevenLabsTTS.ts
import { useState, useCallback } from "react";

type TTSOptions = {
  voiceId?: string; // defaults to a premium voice if omitted
  stability?: number;
  similarityBoost?: number;
};

export function useElevenLabsTTS(
  text: string,
  { voiceId = "EXAMPLE_VOICE_ID", stability = 0.75, similarityBoost = 0.85 }: TTSOptions = {}
) {
  const [loading, setLoading] = useState(false);
  const [error, setError] = useState<string | null>(null);

  const play = useCallback(async () => {
    setLoading(true);
    setError(null);
    try {
      const resp = await fetch(
        `https://api.elevenlabs.io/v1/text-to-speech/${voiceId}`,
        {
          method: "POST",
          headers: {
            "Content-Type": "application/json",
            "xi-api-key": import.meta.env.VITE_ELEVENLABS_API_KEY,
          },
          body: JSON.stringify({
            text,
            model_id: "eleven_monolingual_v1",
            voice_settings: { stability, similarity_boost: similarityBoost },
          }),
        }
      );

      if (!resp.ok) {
        const err = await resp.text();
        throw new Error(`ElevenLabs error: ${err}`);
      }

      const audioBlob = await resp.blob();
      const audioUrl = URL.createObjectURL(audioBlob);
      const audio = new Audio(audioUrl);
      audio.play();
    } catch (e: any) {
      setError(e.message);
    } finally {
      setLoading(false);
    }
  }, [text, voiceId, stability, similarityBoost]);

  return { play, loading, error };
}
Enter fullscreen mode Exit fullscreen mode

Note: The hook uses Vite’s import.meta.env convention. If you’re on Create‑React‑App, replace it with process.env.REACT_APP_ELEVENLABS_API_KEY.


4. Putting It All Together – A Simple UI Component

Now we can build a tiny component that lets users type a sentence and hear it back instantly.

// src/components/VoiceReader.tsx
import { useState } from "react";
import { useElevenLabsTTS } from "../hooks/useElevenLabsTTS";

export function VoiceReader() {
  const [input, setInput] = useState("Hello, world! This is ElevenLabs speaking.");
  const { play, loading, error } = useElevenLabsTTS(input);

  return (
    <div style={{ maxWidth: 600, margin: "2rem auto", padding: "1rem", border: "1px solid #eaeaea", borderRadius: 8 }}>
      <h2>🗣️ Voice‑Enabled UI Demo</h2>

      <textarea
        rows={3}
        style={{ width: "100%", fontSize: "1rem", marginBottom: "0.5rem" }}
        value={input}
        onChange={(e) => setInput(e.target.value)}
      />

      <button
        onClick={play}
        disabled={loading}
        style={{
          padding: "0.5rem 1rem",
          background: loading ? "#999" : "#0070f3",
          color: "#fff",
          border: "none",
          borderRadius: 4,
          cursor: loading ? "default" : "pointer",
        }}
      >
        {loading ? "Generating…" : "Speak"}
      </button>

      {error && <p style={{ color: "red", marginTop: "0.5rem" }}>{error}</p>}
    </div>
  );
}
Enter fullscreen mode Exit fullscreen mode

Drop <VoiceReader /> anywhere in your app (e.g., inside App.tsx). When the user clicks Speak, the hook contacts ElevenLabs, streams back an MP3, and the browser plays it automatically.


5. Going Further – Voice Cloning

One of ElevenLabs’ standout features is voice cloning. If you have a short sample (≈30 seconds) of a speaker, you can upload it via the API and receive a new voice_id that mimics that voice. The workflow is:

  1. POST the audio sample to /v1/voices/add.
  2. Wait for the processing job to finish (the API returns a status URL).
  3. Use the returned voice_id in the TTS request.

Here’s a quick Node script to create a clone:

// cloneVoice.js (run with node)
const fetch = require("node-fetch");
require("dotenv").config();

async function createClone(samplePath) {
  const sample = require("fs").createReadStream(samplePath);
  const resp = await fetch("https://api.elevenlabs.io/v1/voices/add", {
    method: "POST",
    headers: {
      "xi-api-key": process.env.ELEVENLABS_API_KEY,
    },
    body: sample,
  });

  const { voice_id, status_url } = await resp.json();
  console.log(`Clone created! Voice ID: ${voice_id}`);
  console.log(`Check status at: ${status_url}`);
}

createClone("./my-sample.wav");
Enter fullscreen mode Exit fullscreen mode

Once you have the new voice_id, simply pass it to the useElevenLabsTTS hook:

const { play } = useElevenLabsTTS(input, { voiceId: "YOUR_CLONE_ID" });
Enter fullscreen mode Exit fullscreen mode

Caution: Voice cloning should respect privacy and consent. Only clone voices you own or have explicit permission to use.


6. Performance & UX Tips

Issue Fix
Delay before audio plays Use the stream endpoint (/v1/text-to-speech/{voice_id}/stream) to start playback as soon as the first chunk arrives.
Long sentences cause timeouts Break long text into 200‑character chunks and queue them.
Multiple concurrent requests Keep a single Audio instance and cancel any ongoing playback before starting a new one (audio.pause(); audio.currentTime = 0;).
Accessibility Pair the spoken output with aria-live regions so screen readers announce the same content.

7. Deploying to Production

When you move from local dev to production, remember to:

  • Store the API key in a server‑side environment (e.g., Vercel’s Environment Variables).
  • Proxy the request through a serverless function if you want to hide the key entirely from the client.
  • Set appropriate CORS headers – ElevenLabs already allows * for production keys, but a proxy gives you more control.

A minimal Vercel function could look like:

// api/tts.js
export default async function handler(req, res) {
  const { text, voiceId } = req.body;
  const apiResp = await fetch(
    `https://api.elevenlabs.io/v1/text-to-speech/${voiceId}`,
    {
      method: "POST",
      headers: {
        "Content-Type": "application/json",
        "xi-api-key": process.env.ELEVENLABS_API_KEY,
      },
      body: JSON.stringify({
        text,
        model_id: "eleven_monolingual_v1",
      }),
    }
  );

  const audioBlob = await apiResp.blob();
  res.setHeader("Content-Type", "audio/mpeg");
  res.send(Buffer.from(await audioBlob.arrayBuffer()));
}
Enter fullscreen mode Exit fullscreen mode

Your front‑end would now hit /api/tts instead of the raw ElevenLabs endpoint, keeping the key secret.


8. Wrapping Up

Integrating voice AI into a React app is surprisingly straightforward when you have a reliable TTS backend. With just a few lines of code you can:

  • Turn any string into natural‑sounding speech.
  • Swap voices on the fly, even using cloned voices for a personal touch.
  • Keep the UI responsive with proper loading states and error handling.

All of this is powered by ElevenLabs, a service that balances quality, flexibility, and developer ergonomics. If you’ve been curious about adding a voice layer to your next project, now’s the perfect time.


Ready to give your UI a voice? Grab your free ElevenLabs API key, spin up the demo component, and start experimenting. Happy coding! 🚀

Try ElevenLabs today: https://try.elevenlabs.io/kr07zfuqn1bp