Executive Summary
Sermons and teaching sessions are recorded as long audio and video files that are hard to search, summarise or reuse.
A Next.js app that accepts any common audio or video format, extracts a clean audio track with FFmpeg, sends it to Gemini for transcription and analysis, and stores the results. Large uploads fall back to Vercel Blob storage.
Core Capabilities & System Highlights
Any Media Format
Drag-and-drop MP3, WAV, M4A, MP4, MOV and more; audio is extracted automatically.
AI Transcription & Analysis
Gemini's multimodal model produces the transcript and a structured analysis.
Large File Handling
Hybrid upload flow with a Vercel Blob fallback for files too big for a single request.
Saved Results
Transcripts and analyses are stored in MongoDB for later review.
System Architecture & Specs
Next.js App Router
Upload UI and API routes that orchestrate extraction, analysis and storage.
FFmpeg Audio Extraction
fluent-ffmpeg with static FFmpeg binaries produces a clean audio track from any input.
Google Gemini
Multimodal transcription and analysis of the extracted audio.
MongoDB + Vercel Blob
Results in MongoDB via Mongoose; large uploads staged in Vercel Blob.
Key Engineering Trade-offs & Operational Notes
Extracting audio before analysis keeps uploads to the model small and fast.
A hybrid upload flow avoids request-size limits on serverless hosts.
TypeScript types cover the pipeline from upload to stored result.
Tech Stack & Role Matrix
Every library and framework in Faithlence — AI Media Analysis was chosen with intentional architectural trade-offs to balance speed, type safety, security, and scalability.
| Technology | Category | System Role | Performance Rationale |
|---|---|---|---|
| Next.js | Framework | Upload UI and processing API | One app for the interface and the pipeline. |
| FFmpeg | Media | Audio extraction | Handles virtually any audio or video container. |
| Gemini | AI | Transcription and analysis | Native audio understanding in one API call. |
| MongoDB | Database | Result storage | Flexible documents for variable-length analyses. |
Live Preview & Showcase
Metrics & Reliability
Input formats
MP3, WAV, M4A, MP4, MOV and more
Pipeline
Upload, extract, analyse, store
AI model
Multimodal transcription and analysis