Howdy, Stranger!

It looks like you're new here. If you want to get involved, click one of these buttons!


New on LowEndTalk? Please Register and read our Community Rules.

All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.

[Open Source] edgeTTS: Self-hosted Edge TTS server with OpenAI & native streaming APIs

Hi LET guys,

haha, long time busy at working :#

I built a lightweight, self-hosted Edge TTS server called edgeTTS and wanted to share it for who finds it useful for their VPS or homelab.

What it does?

It proxies Microsoft Edge's Read Aloud speech synthesis into a standalone, single-process service with two distinct endpoints:

  1. OpenAI-compatible endpoint (/v1/audio/speech): Works as a drop-in free replacement for third-party tools, readers, translation plugins, and chat UIs (LobeChat, NextChat, STranslate, etc.) that expect standard OpenAI TTS payloads.
  2. Native streaming endpoint (/api/speech): Designed for long documents (up to 20,000 code points) with speed, pitch, and volume adjustments. It splits text losslessly at sentence boundaries and streams the audio back in a single MP3 stream.

Why build it?

Many existing scripts either disconnect on long text or make a brand new TLS/WebSocket handshake for every sentence fragment. edgeTTS reuses upstream WebSocket connections across segments (reducing inter-segment latency down to ~380ms) and includes a bounded concurrency queue (default 4 active streams + 16 FIFO slots) so your server won't spam upstream and trigger IP rate limits.

Resource usage is minimal—the Node process typically idles at around 60–80 MB of RAM, so it happily runs on 512MB / 1GB cheap VPS boxes.

Quick Start (Docker Compose)

services:
  edgetts:
    image: ghcr.io/dejavumoe/edgetts:0.9.5
    container_name: edgetts
    restart: unless-stopped
    read_only: true
    cap_drop:
      - ALL
    security_opt:
      - no-new-privileges:true
    tmpfs:
      - /tmp
    ports:
      - "127.0.0.1:8080:8080"
    environment:
      - NODE_ENV=production
      - API_KEY=replace_with_your_own_secret_key
      - REQUIRE_API_KEY=true

Once running, visit http://127.0.0.1:8080 for the web workbench, or send requests directly to /v1/audio/speech or /api/speech with Bearer auth.

Honest caveats

  • It relies on Microsoft's online Edge service; it's not an offline engine. Audio is piped through memory and never stored on disk.
  • Availability depends on upstream Microsoft endpoints staying accessible from your server's IP.
Sign In or Register to comment.