AKHIL SINGH
PDFVOICE NOTE0:42

A Telegram bot that turns text, PDFs and documents into voice notes, with an optional rewrite into plain conversational speech. The same engine is exposed as a public TTS API and an MCP server.

The problem

I don't enjoy reading long PDFs on a phone. A ten-page paper usually has one paragraph that matters, and I lose interest before I reach it. I wanted to send the thing somewhere and get audio back that I could listen to on a walk.

What I built

A Telegram bot. You send it a line of text or a document (PDF, DOCX, Markdown, CSV, JSON, plain text) and a voice note comes back a few seconds later.

With conversational mode on, it first rewrites the input into how a person would actually say it. A dense paper becomes a few minutes of listenable audio. With the mode off it reads the original word for word, which is useful for proofreading a draft.

Voice, speed and pitch are picked from inline Telegram menus, with previews before you commit. Preferences persist per user.

How it works

One aiohttp webhook service running aiogram. Text messages are buffered briefly so a paste that Telegram splits into several messages is read as one. Documents are downloaded into memory and text is extracted there; message content is never stored.

A voice pipeline does the rest: optional Gemini rewrite, Edge TTS synthesis to MP3, ffmpeg conversion to OGG/Opus so Telegram shows it as a native voice note. MongoDB holds settings, access state and usage counts.

The same synthesis sits behind a public /api/tts endpoint with Swagger docs, a key-protected v2 endpoint with premium voices, and a Streamable HTTP MCP server so AI clients can call it as a tool. Delivery can also go out over the WhatsApp Cloud API.

Outcome

Live as @vaani_tts_bot, deployed with Docker Compose. The API it exposes also powers Whisperian, my Chrome read-aloud extension.

Need something like this built? contact@akhilsingh.in

Next
Whisperian →