Back to the catalog

io.github.huseyinstif/klaket-mcp

Let AI agents watch videos: local transcripts, speakers, scenes, chapters and moment search

Open source Open in the app JSON README (API)

About

Let AI agents watch videos: local transcripts, speakers, scenes, chapters and moment search

Details

Kind
MCP servers
Topic
No topic detected
Publisher
huseyinstif
Origin
official
Category
ferramentas
Transport
local
Version
0.7.1
Stars
2
Last push
2026-07-14T11:55:32Z
Repository state
ativo
Language
Python
License
AGPL-3.0
Added
2026-08-29 04:00:08
Updated
2026-08-29 04:00:08
Origin id
io.github.huseyinstif/klaket-mcp

README

# 🎬 Klaket

**Turn any video into LLM-ready data.**

[![License: AGPL-3.0](https://img.shields.io/badge/license-AGPL--3.0-f5b70f)](LICENSE)
[![PRs welcome](https://img.shields.io/badge/PRs-welcome-4ade80)](CONTRIBUTING.md)
[![Self-host](https://img.shields.io/badge/self--host-docker%20compose%20up-ede8e0)](#quick-start)

![Klaket demo](assets/demo.gif)

> A *klaket* is a clapperboard β€” the tool that syncs sound and image on a film set. **Klaket syncs video with LLMs.**

LLMs read text. The web became readable with scrapers β€” but video, the largest store of human knowledge, is still locked away. Klaket unlocks it: give it a video URL or file, get back structured, timestamped, LLM-ready data.

```bash
pip install klaket
klaket ingest "https://youtube.com/watch?v=..." --wait
```

```jsonc
{
  "transcript": [
    { "start": 14.32, "end": 19.80, "speaker": "S1", "text": "So let's deploy this with docker compose..." }
  ],
  "scenes": [
    { "start": 190.0, "end": 342.5, "keyframes": ["scene_004_01.jpg"] }
  ],
  "chapters": [...],
  "summary": "..."
}
```

## Features

- **πŸ“ Transcript** β€” timestamped speech-to-text in **~100 languages** (auto-detected) with **word-level timestamps**; pick the model per job (`"model": "medium"`)
- **πŸŽ™οΈ Podcasts too** β€” pass an audio file/URL (mp3, m4a…) and Klaket skips the visual stages, deriving chapters from speech pauses
- **πŸ—£οΈ Speaker diarization** β€” who said what (S1/S2/…), local & keyless (sherpa-onnx)
- **πŸ’¬ Subtitles** β€” ready-to-use `.srt` / `.vtt` files with speaker labels
- **🎞️ Scene detection** β€” content-aware scene boundaries + keyframes per scene
- **πŸ”Ž On-screen text (OCR)** β€” reads slides, terminals and captions per scene, local & keyless
- **🧩 One JSON timeline** β€” transcript, scenes, frames and on-screen text aligned on a single timeline
- **πŸ”Œ Works offline, no API key required** β€” the core pipeline uses zero LLM calls
- **🧠 Pluggable model layer** β€” optional scene descriptions via local VLMs (Ollama) or any OpenAI-compatible endpoint (`KLAKET_VLM=off` by default)
- **πŸ€– MCP server** β€” let coding agents "watch" any video and find moments inside it
- **πŸ” In-video search** β€” `GET /v1/jobs/{id}/search?q=…` finds the exact moment
- **▢️ Playground** β€” the dashboard plays the video with a click-to-seek, live-highlighted transcript

## SDKs

```python
# pip install klaket
from klaket import Klaket
result = Klaket().process("https://youtube.com/watch?v=...", num_speakers=2)
```

```ts
// npm i klaket-sdk
import { Klaket } from "klaket-sdk";
const result = await new Klaket().process("https://youtube.com/watch?v=...");
```

## Give your agent eyes

```bash
# Claude Code
claude mcp add klaket -- npx klaket-mcp   # KLAKET_API_URL defaults to localhost:8484
```

Then: *"Watch https://youtube.com/watch?v=… and summarize the commands the presenter runs."*
The agent gets `klaket_ingest`, `klaket_job_status` and `klaket_get_result` tools.

## Quick start

```bash
git clone https://github.com/huseyinstif/klaket.git && cd klaket
docker compose up --build
# API on :8484, dashboard on :5180
curl -X POST localhost:8484/v1/ingest \
  -H "Content-Type: application/json" \
  -d '{"url": "https://youtube.com/watch?v=..."}'
```

That's it β€” no API keys, no GPUs required. `make help` lists developer shortcuts (`make up`, `make test`, `make e2e`).

## Architecture

```
client ──► Go API ──► Redis queue ──► Python worker (ffmpeg Β· faster-whisper Β· scenedetect)
                β”‚                          β”‚
            dashboard β—„β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   /data/jobs/<id>/result.json
```

- `apps/api` β€” Go, job orchestration
- `apps/worker` β€” Python, media pipeline
- `apps/dashboard` β€” React dashboard

## Self-host vs Cloud

Klaket is open source (AGPL-3.0) and fully self-hostable. A hosted, pay-per-minute cloud API with managed GPUs is planned β€” join the waitlist (coming soon).

## Status

🚧 v0.7 β€” pre-1.0, moving fast. Star the repo to follow along.

## License

[AGPL-3.0](LICENSE). SDKs and clients will be MIT.

## Contact

Built by HΓΌseyin TΔ±ntaş β€” [X (@1337stif)](https://x.com/1337stif) Β· [LinkedIn](https://www.linkedin.com/in/huseyintintas/)

More