Tutorials & Guides
How to Build a Telegram Bot to Transcribe and Summarise Voice Notes Using Gemini 1.5 Flash and Python
Ditch the agony of listening to four-minute rambling audio messages. Build a lightweight Python bot that feeds voice notes directly to Gemini Flash for instant, formatted summaries.
Updated 9/7/2026
We have all been there. You are in a meeting, or perhaps enjoying a quiet pint, and your phone buzzes with a voice note. You look at the progress bar: four minutes and thirty-seven seconds of unstructured, rambling consciousness. Listening to it is out of the question, and typing 'can you just text me?' makes you look like a misanthrope.
Instead of suffering, we are going to build a solution. In this guide, we will write a lightweight Python Telegram bot that intercepts voice notes, feeds them directly to Google's Gemini 1.5 Flash, and spits back a clean, bulleted summary.
Why Gemini 1.5 Flash? Because unlike most LLMs that require you to run your audio through a separate speech-to-text API (like Whisper) before feeding the text to the model, Gemini 1.5 Flash accepts audio natively. It is incredibly fast, shockingly cheap, and handles nuances in spoken tone like a charm.
Why Native Audio Processing Matters
When you use a multi-step pipeline (Audio -> Whisper -> LLM), you lose a lot of context. Verbal emphasis, pauses, and emotional tone get flattened into plain text. Worse, if your speech-to-text engine mishears a critical industry term or acronym, the downstream LLM will confidently hallucinate a summary based on that error.
By using /platforms/gemini, the model listens to the audio file directly. It bypasses the middleman, preserves the speaker’s intent, and keeps your latency remarkably low. If you run into issues with audio limits or rate limits during setup, you can check the Gemini support pages for regional API quotas.
Prerequisites
Before we begin writing code, you will need:
Python 3.10+* installed on your local machine or server.
A Telegram Bot Token*. You can get this by messaging the @BotFather on Telegram and typing /newbot.
A Gemini API Key*, which you can grab from Google AI Studio.
Step 1: Setting Up the Environment
First, let’s create a virtual environment and install our dependencies. We will need python-telegram-bot to interface with Telegram, google-generativeai to talk to Gemini, and pydub to handle audio conversion (since Telegram voice notes are sent as .ogg files and Gemini prefers standard formats like .wav or .mp3).
`bash
mkdir voice-summariser-bot
cd voice-summariser-bot
python3 -m venv venv
source venv/bin/activate
pip install python-telegram-bot google-generativeai pydub
`
Note: pydub requires ffmpeg to process audio files. If you don't have it installed on your system, grab it via your package manager:
macOS*: brew install ffmpeg
Ubuntu/Debian*: sudo apt install ffmpeg
Step 2: Writing the Audio Converter
Telegram encodes voice notes using the OGG/Opus codec. While Gemini can sometimes handle raw OGG, converting it to a standard, uncompressed WAV file ensures we never run into decoding errors at the API level.
Let’s create a file called audio_utils.py to handle this logic cleanly:
`python
import os
from pydub import AudioSegment
def convert_ogg_to_wav(ogg_path: str, wav_path: str) -> bool:
try:
audio = AudioSegment.from_file(ogg_path, format="ogg")
audio.export(wav_path, format="wav")
return True
except Exception as e:
print(f"Error converting audio: {e}")
return False
`
Step 3: Setting Up the Telegram Bot and Gemini Orchestrator
Now, let's build the core application in bot.py. We will initialise the Telegram bot, intercept any incoming voice messages, download them locally, convert them to WAV, and push them to Gemini.
What makes this whole setup tick is our system prompt. We need to tell Gemini exactly how to structure the output so we can scan the summary in under three seconds.
`python
import os
import logging
from telegram import Update
from telegram.ext import Application, MessageHandler, filters, ContextTypes
import google.generativeai as genai
from audio_utils import convert_ogg_to_wav
Configure logging logging.basicConfig( format="%(asctime)s - %(name)s - %(levelname)s - %(message)s", level=logging.INFO )
Grab keys from environmental variables TELEGRAM_TOKEN = os.environ.get("TELEGRAM_TOKEN") GEMINI_API_KEY = os.environ.get("GEMINI_API_KEY")
if not TELEGRAM_TOKEN or not GEMINI_API_KEY: raise ValueError("Please set TELEGRAM_TOKEN and GEMINI_API_KEY env variables.")
genai.configure(api_key=GEMINI_API_KEY)
async def handle_voice(update: Update, context: ContextTypes.DEFAULT_TYPE): # 1. Inform the user we are on it status_message = await update.message.reply_text("📥 Downloading your monologue... ") # Create paths voice_file = await update.message.voice.get_file() ogg_path = f"voice_{update.message.message_id}.ogg" wav_path = f"voice_{update.message.message_id}.wav"
try: # 2. Download the voice note await voice_file.download_to_drive(ogg_path) await status_message.edit_text("⚡ Converting audio...") # 3. Convert to WAV if not convert_ogg_to_wav(ogg_path, wav_path): await status_message.edit_text("❌ Failed to process audio formatting.") return await status_message.edit_text("🧠 Gemini is listening...") # 4. Upload file to Gemini File API # Using the File API is recommended for audio processing audio_file = genai.upload_file(path=wav_path) # 5. Build our prompt. You can generate custom variations using our /prompts generator. system_prompt = ( "You are an elite executive assistant. Your job is to listen to the provided voice note " "and write a ruthless, ultra-concise summary. " "Structure your response as follows:\n" "TL;DR: (A one-sentence summary of the main point)\n\n" "Key Takeaways:\n- (Bullet points of actionable items, decisions, or core messages)\n\n" "Tone/Urgency: (E.g., Casual but urgent, low-priority rambling, frustrated, etc.)" ) model = genai.GenerativeModel("gemini-1.5-flash") response = model.generate_content([audio_file, system_prompt]) # 6. Clean up files on Gemini and locally genai.delete_file(audio_file.name) # 7. Reply to user with the summary await status_message.edit_text(response.text, parse_mode="Markdown")
except Exception as e: logging.error(f"Error handling voice note: {e}") await status_message.edit_text("💀 Something went sideways while processing that audio.") finally: # Clean up local files for path in [ogg_path, wav_path]: if os.path.exists(path): os.remove(path)
def main(): application = Application.builder().token(TELEGRAM_TOKEN).build() # Register handler for voice messages application.add_handler(MessageHandler(filters.VOICE, handle_voice)) logging.info("Bot started. Listening for voice notes...") application.run_polling()
if __name__ == "__main__":
main()
`
Step 4: Tweaking Your Prompt for Personality
If you find the summaries a little too dry, modify the system prompt to match your personal context. For example, if you get a lot of developer feedback over voice notes, tweak the prompt to focus specifically on bug reports, reproduction steps, and system architecture. Feel free to use our /prompts system to test and optimise your instructions before committing them to your bot.
Running Your Bot
To spin up your bot locally, simply run:
`bash
export TELEGRAM_TOKEN="your-telegram-token"
export GEMINI_API_KEY="your-gemini-key"
python bot.py
`
Forward a voice note to your new bot, and watch the magic happen. Within a couple of seconds, you'll have a structured, skimmable breakdown of an otherwise tedious five-minute recording. No more awkward pauses, no more conversational fillers, just pure raw data.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.