Skip to main content
Glama
Knuckles-Team

audio-transcriber

Audio Transcriber

CLI or API | MCP | Agent

PyPI - Version MCP Server PyPI - Downloads GitHub Repo stars GitHub forks GitHub contributors PyPI - License GitHub GitHub last commit (by committer) GitHub pull requests GitHub closed pull requests GitHub issues GitHub top language GitHub language count GitHub repo size GitHub repo file count (file type) PyPI - Wheel PyPI - Implementation

Version: 2.0.0

Documentation — Installation, deployment, and usage across the CLI, Python API, MCP server, and A2A agent are maintained in the official documentation.


Related MCP server: audio-transcription-mcp

Overview

Audio Transcriber is a production-grade Agent and Model Context Protocol (MCP) server designed to interface directly with Transcribe your .wav .mp4 .mp3 .flac files to text or record your own audio!.


Key Features

  • Consolidated Action-Routed MCP Tools: Minimizes token overhead and eliminates tool bloat in LLM contexts by grouping methods into optimized, togglable tool modules.

  • Enterprise-Grade Security: Comprehensive support for Eunomia policies, OIDC token delegation, and granular execution context tracking.

  • Integrated Graph Agent: Built-in Pydantic AI agent supporting the Agent Control Protocol (ACP) and standard Web interfaces (AG-UI).

  • Native Telemetry & Tracing: Out-of-the-box OpenTelemetry exports and native Langfuse tracing.


CLI or API

This agent wraps the Transcribe your .wav .mp4 .mp3 .flac files to text or record your own audio! API. You can interact with it programmatically or via its integrated execution entrypoints.

Detailed instructions on how to use the underlying API wrappers, extended schema bindings, and developer SDK references are maintained in docs/index.md.


MCP

This server utilizes dynamic Action-Routed tools to optimize token overhead and maximize IDE compatibility.

Available MCP Tools

The table below is auto-generated from the live server — do not edit by hand.

Condensed action-routed tools (default — MCP_TOOL_MODE=condensed)

MCP Tool

Toggle Env Var

Description

health_check

MISCTOOL

transcribe_audio

AUDIO_PROCESSINGTOOL

Transcribes audio from a provided file or by recording from the microphone.

Verbose 1:1 API-mapped tools (MCP_TOOL_MODE=verbose or both)

MCP Tool

Toggle Env Var

Description

audio_transcriber_export

AUDIO_TRANSCRIBERTOOL

Export transcription to specified formats.

audio_transcriber_initiate_stream

AUDIO_TRANSCRIBERTOOL

Initiate the audio input stream.

audio_transcriber_interact

AUDIO_TRANSCRIBERTOOL

Interact with PersonaPlex server via WebSocket.

audio_transcriber_record

AUDIO_TRANSCRIBERTOOL

Record audio for a specified duration or until stopped.

audio_transcriber_save_stream

AUDIO_TRANSCRIBERTOOL

Save the recorded frames to a WAV file.

audio_transcriber_stop_stream

AUDIO_TRANSCRIBERTOOL

Stop and close the audio stream.

audio_transcriber_transcribe

AUDIO_TRANSCRIBERTOOL

Transcribe the audio file using the initialized backend.

2 action-routed tool(s) (default) · 7 verbose 1:1 tool(s). Each is enabled unless its <DOMAIN>TOOL toggle is set false; MCP_TOOL_MODE selects the surface (condensed default · verbose 1:1 · both). Auto-generated — do not edit.

Detailed tool schemas, parameter shapes, and validation constraints are preserved in docs/usage.md.

Dynamic Tool Selection & Visibility

This MCP server supports dynamic toolset selection and visibility filtering at runtime. This allows you to restrict the set of exposed tools in order to prevent blowing up the LLM's context window.

You can configure tool filtering via multiple input channels:

  • CLI Arguments: Pass --tools or --toolsets (or their disabled counterparts --disabled-tools and --disabled-toolsets) during startup.

  • Environment Variables: Define standard environment variables:

    • MCP_ENABLED_TOOLS / MCP_DISABLED_TOOLS

    • MCP_ENABLED_TAGS / MCP_DISABLED_TAGS

  • HTTP SSE Request Headers: Pass custom headers during transport initialization:

    • x-mcp-enabled-tools / x-mcp-disabled-tools

    • x-mcp-enabled-tags / x-mcp-disabled-tags

  • HTTP SSE Request Query Parameters: Append query parameters directly to your transport connection URL:

    • ?tools=tool1,tool2

    • ?tags=tag1

When query strings or parameters are supplied, an LLM-free Knowledge Graph resolution layer (using DynamicToolOrchestrator) matches query intents against known tool tags, names, or descriptions, with safe fallback and automated 24-hour background cache refreshing.


MCP Configuration Examples

Install the connector-focused [mcp] extra. Examples use audio-transcriber[mcp] to add FastMCP / FastAPI through agent-utilities[mcp]; the required Agent Utilities core still carries epistemic-graph[full]. The [agent-runtime] extra additionally enables model orchestration.

stdio Transport (local IDEs — Cursor, Claude Desktop, VS Code)

{
  "mcpServers": {
    "audio-transcriber-mcp": {
      "command": "uvx",
      "args": [
        "--from",
        "audio-transcriber[mcp]",
        "audio-transcriber-mcp"
      ],
      "env": {
        "MCP_TOOL_MODE": "intent",
        "AUDIO_PROCESSINGTOOL": "True",
        "MISCTOOL": "True",
        "TRANSCRIBE_DIRECTORY": "/path/to/transcribe_directory",
        "WHISPER_MODEL": "base"
      }
    }
  }
}

Runtime references require an alias-aware launcher such as GraphOS. Other launchers must omit those entries and inject the resolved values through their own runtime secret boundary.

Streamable-HTTP Transport (networked / production)

{
  "mcpServers": {
    "audio-transcriber-mcp": {
      "command": "uvx",
      "args": [
        "--from",
        "audio-transcriber[mcp]",
        "audio-transcriber-mcp",
        "--transport",
        "streamable-http",
        "--port",
        "8000"
      ],
      "env": {
        "TRANSPORT": "streamable-http",
        "HOST": "127.0.0.1",
        "PORT": "8000",
        "MCP_TOOL_MODE": "intent",
        "AUDIO_PROCESSINGTOOL": "True",
        "MISCTOOL": "True",
        "TRANSCRIBE_DIRECTORY": "/path/to/transcribe_directory",
        "WHISPER_MODEL": "base"
      }
    }
  }
}

Alternatively, connect to a pre-deployed Streamable-HTTP instance by url:

{
  "mcpServers": {
    "audio-transcriber-mcp": {
      "url": "http://localhost:8000/audio-transcriber-mcp/mcp"
    }
  }
}

Run a reviewed container image as a least-privilege stdio child (no listener or published port):

docker run -i --rm \
  --read-only \
  --cap-drop=ALL \
  --security-opt=no-new-privileges \
  --pids-limit=256 \
  --tmpfs /tmp:rw,noexec,nosuid,nodev,size=64m \
  -e TRANSPORT=stdio \
  -e MCP_TOOL_MODE=intent \
  -e AUDIO_PROCESSINGTOOL=True \
  -e MISCTOOL=True \
  -e TRANSCRIBE_DIRECTORY=/path/to/transcribe_directory \
  -e WHISPER_MODEL=base \
  registry.example.invalid/audio-transcriber@sha256:<digest> audio-transcriber-mcp

For containerized network HTTP, supply an authenticated TLS ingress (or direct server TLS), exact MCP_ALLOWED_HOSTS, and an exact trusted-proxy CIDR policy through the operator-owned deployment profile. The generator does not emit an unauthenticated non-loopback listener.

Auto-generated from the code-read env surface (MCP_TOOL_MODE + package vars) — do not edit.

Additional Deployment Options

audio-transcriber can run as a local stdio process or container, or behind a remote network boundary. The Deployment guide carries the detailed transport contract.

  • Local container — launch a reviewed immutable image as a least-privilege stdio child with no listener or published port.

  • Remote URL — connect through an operator-supplied authenticated HTTPS ingress. Keep its URL, outbound identity references, trust profile, and exact MCP_ALLOWED_HOSTS in AgentConfig.

Agent

This repository features a fully integrated Pydantic AI Graph Agent. It communicates over the Agent Control Protocol (ACP) and interacts seamlessly with the Agent Web UI (AG-UI) and Terminal interface.

Running the Agent CLI

To start the interactive command-line agent:

# Configure transcription (optional)
export WHISPER_MODEL="base"
export TRANSCRIBE_DIRECTORY="/path/to/transcribe_directory"

# Run the agent server
audio-transcriber-agent --provider openai --model-id gpt-4o

Docker Compose Orchestration

The following docker/agent.compose.yml configures the Agent, Web UI, and Terminal Interface together:

version: '3.8'

services:
  audio-transcriber-mcp:
    image: example/audio-transcriber:mcp
    container_name: audio-transcriber-mcp
    hostname: audio-transcriber-mcp
    restart: always
    env_file:
      - ../.env
    environment:
      - PYTHONUNBUFFERED=1
      - HOST=0.0.0.0
      - PORT=8000
      - TRANSPORT=streamable-http
    ports:
      - "8000:8000"
    healthcheck:
      test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/health')"]
      interval: 30s
      timeout: 10s
      retries: 3
      start_period: 10s
    logging:
      driver: json-file
      options:
        max-size: "10m"
        max-file: "3"

  audio-transcriber-agent:
    image: example/audio-transcriber@sha256:<digest>
    container_name: audio-transcriber-agent
    hostname: audio-transcriber-agent
    restart: always
    depends_on:
      - audio-transcriber-mcp
    env_file:
      - ../.env
    command: [ "audio-transcriber-agent" ]
    environment:
      - PYTHONUNBUFFERED=1
      - HOST=0.0.0.0
      - PORT=9014
      - MCP_URL=http://audio-transcriber-mcp:8000/mcp
      - PROVIDER=${PROVIDER:-openai}
      - MODEL_ID=${MODEL_ID:-gpt-4o}
      - ENABLE_WEB_UI=True
      - ENABLE_OTEL=True
    ports:
      - "9014:9014"
    healthcheck:
      test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:9014/health')"]
      interval: 30s
      timeout: 10s
      retries: 3
      start_period: 10s
    logging:
      driver: json-file
      options:
        max-size: "10m"
        max-file: "3"

Detailed graph node architecture explanations, custom skill configurations, and agentic trace guides are available in docs/deployment.md.


Security & Governance

Built directly upon the enterprise-ready agent-utilities core, standard security parameters are fully supported:

Access Control & Policy Enforcement

  • Eunomia Policies: Fine-grained, policy-driven tool authorization. Supports none, local embedded (mcp_policies.json), or centralized remote modes.

  • OIDC Token Delegation: Compliant with RFC 8693 token exchange for flowing authenticating user credentials from Web UI / ACP → Agent → MCP.

  • Scoped Credentials: Execution context runs restricted to the specific caller identity.

Runtime Security Grid

Feature

Functionality

Enablement

Tool Guard

Sensitivity inspection with human-in-the-loop validation

Enabled by default

Prompt Injection Defense

Input scanning, repetition monitoring, and recursive loop blocks

Enabled by default

Context Safety Guard

Stuck-loop detectors and contextual overflow preemptive alerts

Enabled by default


Environment Variables

Package environment variables

Variable

Example

Description

HOST

0.0.0.0

PORT

8000

TRANSPORT

stdio

options: stdio, streamable-http, sse

ENABLE_OTEL

True

OTEL_EXPORTER_OTLP_ENDPOINT

http://localhost:8080/api/public/otel

OTEL_EXPORTER_OTLP_PUBLIC_KEY

pk-...

OTEL_EXPORTER_OTLP_SECRET_KEY

sk-...

OTEL_EXPORTER_OTLP_PROTOCOL

http/protobuf

EUNOMIA_TYPE

none

options: none, embedded, remote

EUNOMIA_POLICY_FILE

mcp_policies.json

EUNOMIA_REMOTE_URL

http://eunomia-server:8000

TRANSCRIBE_DIRECTORY

/path/to/transcribe_directory

Directory where transcripts are written (defaults to the data dir under audio-transcriber)

MISCTOOL

True

AUDIO_PROCESSINGTOOL

True

WHISPER_MODEL

base

Standard OpenAI Whisper model to use for local transcription (e.g., base, tiny, small)

Inherited agent-utilities variables (apply to every connector)

Variable

Example

Description

MCP_TOOL_MODE

condensed

Tool surface: condensed

MCP_ENABLED_TOOLS

Comma-separated tool allow-list

MCP_DISABLED_TOOLS

Comma-separated tool deny-list

MCP_ENABLED_TAGS

Comma-separated tag allow-list

MCP_DISABLED_TAGS

Comma-separated tag deny-list

MCP_CLIENT_AUTH

Outbound MCP auth (oidc-client-credentials for fleet calls)

OIDC_CLIENT_ID

OIDC client id (service-account auth)

OIDC_CLIENT_SECRET

OIDC client secret (service-account auth)

DEBUG

False

Verbose logging

PYTHONUNBUFFERED

1

Unbuffered stdout (recommended in containers)

MCP_URL

http://localhost:8000/mcp

URL of the MCP server the agent connects to

PROVIDER

openai

LLM provider for the agent

MODEL_ID

gpt-4o

Model id for the agent

ENABLE_WEB_UI

True

Serve the AG-UI web interface

15 package + 14 inherited variable(s). Auto-generated from .env.example + the shared agent-utilities set — do not edit.

Every variable the server reads, grouped by purpose.

Transcription

Variable

Description

Default

WHISPER_MODEL

Local OpenAI Whisper model (e.g. base, tiny, small)

base

TRANSCRIBE_DIRECTORY

Directory where transcripts are written

data dir

MCP server / transport

Variable

Description

Default

TRANSPORT

stdio, streamable-http, or sse

stdio

HOST

Bind host (HTTP transports)

0.0.0.0

PORT

Bind port (HTTP transports)

8000

MCP_TOOL_MODE

Tool surface: condensed, verbose, or both

condensed

MCP_ENABLED_TOOLS / MCP_DISABLED_TOOLS

Comma-separated tool allow/deny list

MCP_ENABLED_TAGS / MCP_DISABLED_TAGS

Comma-separated tag allow/deny list

DEBUG

Verbose logging

False

PYTHONUNBUFFERED

Unbuffered stdout (recommended in containers)

1

Tool toggles

Each action-routed tool can be disabled individually via its toggle env var (set to false). See the Available MCP Tools table above for the authoritative names.

Variable

Description

Default

MISCTOOL

Toggle the miscellaneous / health-check tool

True

AUDIO_PROCESSINGTOOL

Toggle the audio-processing (transcription) tool

True

Telemetry & governance

Variable

Description

Default

ENABLE_OTEL

Enable OpenTelemetry export

True

OTEL_EXPORTER_OTLP_ENDPOINT

OTLP collector endpoint

OTEL_EXPORTER_OTLP_PUBLIC_KEY / OTEL_EXPORTER_OTLP_SECRET_KEY

OTLP auth keys

OTEL_EXPORTER_OTLP_PROTOCOL

OTLP protocol (e.g. http/protobuf)

EUNOMIA_TYPE

Authorization mode: none, embedded, remote

none

EUNOMIA_POLICY_FILE

Embedded policy file

mcp_policies.json

EUNOMIA_REMOTE_URL

Remote Eunomia server URL

Agent CLI (full [agent] runtime only)

Variable

Description

Default

MCP_URL

URL of the MCP server the agent connects to

http://localhost:8000/mcp

PROVIDER

LLM provider (e.g. openai)

openai

MODEL_ID

Model id (e.g. gpt-4o)

gpt-4o

ENABLE_WEB_UI

Serve the AG-UI web interface

True

See .env.example for a copy-paste starting point.


Installation

Pick the extra that matches what you want to run:

Extra

Installs

Use when

audio-transcriber[mcp]

Connector-focused MCP server (agent-utilities[mcp] — FastMCP/FastAPI + epistemic-graph[full])

You only run the MCP server (smallest install / image)

audio-transcriber[agent]

Agent runtime (agent-utilities[agent-runtime,logfire] — model orchestration + epistemic-graph[full])

You run the integrated agent

audio-transcriber[all]

Everything (mcp + agent)

Development / both surfaces

# Connector-focused MCP server (includes the shared graph engine)
uv pip install "audio-transcriber[mcp]"

# Agent runtime (adds model orchestration to the shared graph engine)
uv pip install "audio-transcriber[agent]"

# Everything (development)
uv pip install "audio-transcriber[all]"      # or: python -m pip install "audio-transcriber[all]"

Container images (:mcp vs :agent)

One multi-stage docker/Dockerfile builds two right-sized images, selected by --target:

Image tag

Build target

Contents

Entrypoint

example/audio-transcriber:mcp

--target mcp

audio-transcriber[mcp]connector-focused, includes epistemic-graph[full]; no model-orchestration stack

audio-transcriber-mcp

example/audio-transcriber@sha256:<digest>

--target agent (default)

audio-transcriber[agent]agent runtime, model orchestration + epistemic-graph[full]

audio-transcriber-agent

docker build --target mcp   -t example/audio-transcriber:mcp    docker/   # connector-focused MCP server
docker build --target agent -t example/audio-transcriber:agent-local docker/   # agent runtime

docker/mcp.compose.yml runs the connector-focused :mcp server; docker/agent.compose.yml runs the agent (immutable agent digest) with a co-located :mcp sidecar.

Knowledge-graph database (epistemic-graph)

Both [mcp] and [agent] carry the epistemic-graph engine through the required Agent Utilities core dependency (epistemic-graph[full]). The [mcp] extra keeps the server connector-focused; [agent] additionally enables model orchestration. Local deployments can use the bundled engine. For production or shared state, run epistemic-graph as a dedicated database service and configure the runtime to use it. Deployment recipes (single-node + Raft HA), connection configuration, and architecture diagrams are documented in the epistemic-graph deployment guide.


Repository Owners

GitHub followers GitHub User's stars


Documentation

The complete documentation is published as the official documentation site and is the recommended reference for installation, deployment, and day-to-day operation.

Page

Contents

Installation

pip, source, extras, prebuilt Docker image

Deployment

run the MCP server and agent, Compose, Caddy + Technitium, env config

Usage

the MCP tool, the AudioTranscriber API, the CLI

Overview

capability summary and ecosystem role

Concepts

concept registry (CONCEPT:AUDIO-*)


Contribute

Contributions are welcome! Please ensure code quality by executing local checks before submitting pull requests:

  • Format code using ruff format .

  • Lint code using ruff check .

  • Validate type-safety with mypy .

  • Execute test suites using pytest

Deploy with agent-utilities-deployment

Provision this package with the consolidated agent-utilities-deployment workflow. It selects an installed-package, editable-source, or immutable-container path; records only runtime secret and TLS-profile references in AgentConfig; and runs doctor, registration, policy, observability, and rollback gates. Ask your agent to "deploy audio-transcriber with agent-utilities-deployment".

Install mode

Command

Installed package

uv tool install "audio-transcriber[mcp]", then run audio-transcriber-mcp

Editable source

uv pip install -e ".[agent]", then run audio-transcriber-mcp

Immutable container

deploy registry.example.invalid/audio-transcriber@sha256:<digest> through the operator-selected orchestrator

The repository embeds no deployment profile, credential value, certificate path, or environment-specific endpoint. Supply those at runtime through AgentConfig and the configured secret provider.

Governed capability contract

This package ships a compact canonical skill surface with specialist procedures kept as referenced workflows. The current MCP tools, skill metadata, connector_manifest.yml, ontology, mappings, shapes, fixtures, migrations, tool-schema fingerprints, and certification metadata form one versioned capability contract. Validate them together; do not rely on stale tool names or historical per-task skill wrappers.

Runtime endpoints, credentials, certificate trust, tenant identity, retention, and observability policy are deployment inputs and are never packaged values. See Configuration, trust, and privacy before enabling a network transport, connector ingestion, GraphOS delegation, or trace export.

A
license - permissive license
-
quality - not tested
B
maintenance

Maintenance

Maintainers
Response time
3dRelease cycle
94Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

View all related MCP servers

Related MCP Connectors

  • OCR, transcription, file extraction, and image generation for AI agents via MCP.

  • MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)

  • Official MCP server for OmniDimension. Drive voice agents, dispatch calls, and run bulk campaigns.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Knuckles-Team/audio-transcriber'

If you have feedback or need assistance with the MCP directory API, please join our Discord server