An autonomous, brand-aware video studio for Descript, driven entirely in plain language.
Give it a prompt or drop in footage. It loads your cached brand profile, shows a short build plan, then builds the whole video end-to-end — cleanup, demo polish, brand kit, captions — verifying every step. After that you just talk to it: "make the captions bigger," "slow the intro," "give me a vertical cut," "publish it."
You: make me a demo video from this screen recording, with captions
Plan: Source = "Beta Launch - Take 1". Cleanup (filler + silences)
→ demo polish (trim dead time, light zooms) → brand intro/outro
→ captions → ~2 min, 16:9. Reply "go" to build.
You: go
✓ imported 0:00 → 3:41
✓ cleanup −1:12 (filler + silences)
✓ polish −0:47 (dead time trimmed, 4 zooms)
✓ brand intro 3s, outro 4s, lower-third
✓ captions brand-styled, lower third
Done — "Beta Launch - Final", 1:42, 16:9.
Want a vertical version, or should I publish?
- What this is · Requirements · Install
- How it works · Commands · Skills · Repo map
- The subagent · Configuration · What it will not do
- Development · Documentation · Known limitations · License
A Claude Code / Cowork plugin — a bundle of markdown skills, one subagent, and shared reference docs. There is no compiled code and no build step. Claude reads these files and drives the Descript API through an MCP connector.
It is built around three ideas:
- One plan, then autonomy. You approve a short build plan once. The
descript-editorsubagent then runs every pass back-to-back without stopping to ask permission at each step. - Brand is cached, not re-asked. A one-time brand setup resolves colors, fonts, logo, caption style, and CTA into
~/.descript-studio/brand-profile.md. Every brand-aware skill reads it. You never re-answer "what's your brand font?" - Every job is verified. Descript edits are async and not cleanly undoable, so each pass follows
resolve → validate → act → wait → verifyand diffs the result before continuing. See resilience.md.
| Connector | Required? | Used for |
|---|---|---|
| Descript | Yes | Almost everything. 11 of the 12 skills drive Descript through it (import_media, prompt_project_agent, get_project, publish_project, …). Only descript-brand-setup runs entirely locally. |
| Figma | No | Auto-filling brand colors/fonts/logo during brand setup |
| brand-voice | No | Auto-filling tone, do/don't, and CTA during brand setup |
The plugin ships a working .mcp.json pointing at Descript's official remote MCP server:
{
"mcpServers": {
"descript": {
"type": "http",
"url": "https://api.descript.com/v2/mcp"
}
}
}Authentication is OAuth — no API token. The URL is the same for every account; your login during the OAuth flow is what ties the connection to your Descript Drive. In Cowork, the host-provided Descript connector is used instead and this declaration simply records the dependency.
Note
MCP usage draws on your Descript plan: imports consume media minutes, and Underlord edits (Studio Sound, filler removal, captions) consume AI credits.
Without Figma or brand-voice, brand setup falls back to a short guided fill — one round of questions, then it caches the answers.
Option A — marketplace (recommended):
/plugin marketplace add katekruger/descript-studio
/plugin install descript-studio@descript-studioOption B — clone the source, if you want to read or modify the skills:
git clone https://github.com/katekruger/descript-studio.git
claude plugin install .Option C — packaged bundle: download dist/descript-studio.plugin and install it directly. See dist/README.md for what each artifact is.
First run will have no brand profile, so it runs descript-brand-setup once before building. Say "refresh brand" any time to update it.
┌──────────────────┐
your prompt ───► │ descript-studio │ the brain: loads brand, resolves
or footage │ (orchestrator) │ source, proposes a build plan
└────────┬─────────┘
│ you reply "go"
▼
┌──────────────────┐
│ descript-editor │ the subagent: runs passes
│ (autonomous) │ back-to-back, verifies each
└────────┬─────────┘
▼
import ──► cleanup ──► demo polish ──► brand kit ──► captions ──► review
│
▼ then, conversationally:
descript-edit · descript-export · descript-clips · descript-broll
The skills that drive Descript obey the same shared references, so behavior stays consistent and a fix in one place propagates everywhere. Not every skill needs all four — descript-brand-setup makes no Descript calls and reads only the brand schema, and descript-review is read-only:
| Reference | What it governs |
|---|---|
prompt-patterns.md |
The exact phrasings sent to prompt_project_agent. Single source of truth — improve one pattern, every skill benefits. |
resilience.md |
The edit loop, validation, retry rules, error translation, autonomous vs. checkpoint mode. |
responsiveness.md |
Time estimates per operation and the "never go silent" rule during long jobs. |
brand-profile.md |
The brand profile schema and cache-resolution rules. |
Nine slash commands cover the main surfaces. Every one takes optional arguments, and each falls back to asking if you give it none.
| Command | Does | Example |
|---|---|---|
/video-new |
Plan and build a video end to end | /video-new a 2-min demo from this recording, with captions |
/video-brand |
Set up or refresh the cached brand profile | /video-brand |
/video-cleanup |
Filler words, silences, retakes, Studio Sound | /video-cleanup |
/video-polish |
Screen-demo polish: dead time, zooms | /video-polish focus on the checkout flow |
/video-captions |
Brand-styled captions | /video-captions large, centered for social |
/video-edit |
One plain-language change | /video-edit make the captions bigger |
/video-clips |
Find highlights, cut vertical shorts | /video-clips |
/video-publish |
Publish, export transcript, or reframe | /video-publish 16:9 and 9:16 |
/video-status |
What's in the project, what changed | /video-status |
Natural language works everywhere too — the commands are a shortcut, not a requirement. "Clean up my recording" and /video-cleanup reach the same skill.
12 skills plus one subagent. Full detail — triggers, MCP calls, pipeline position, guardrails — in docs/skills.md.
| Skill | Does |
|---|---|
descript-studio |
Orchestrator. Plan → autonomous build → route follow-ups |
descript-brand-setup |
Resolve and cache the brand profile (one-time, refreshable) |
descript-import |
Create a project, upload/import media, validate before uploading |
descript-cleanup |
Filler words, silences, retakes, Studio Sound |
descript-demo-polish |
Screen demos: trim dead time, speed slow stretches, motivated zooms |
descript-brandkit |
Intro/outro, lower-thirds, brand fonts and colors |
descript-captions |
Brand-styled captions; large word-by-word for social |
descript-broll |
Stock footage at described moments; cover jarring cuts |
descript-edit |
The iteration surface. Any plain-language tweak, scoped tightly |
descript-export |
Publish, export transcript, reframe to 9:16 / 1:1 |
descript-clips |
Find highlights, cut vertical social shorts |
descript-review |
Read-only status: what's in the project, what changed |
descript-editor (subagent) |
Autonomous executor for an approved build plan |
.claude-plugin/plugin.json plugin manifest (name, version, author, agents path)
.mcp.json Descript MCP server config (official remote endpoint, OAuth)
commands/video-*.md 9 slash commands, thin wrappers that invoke a skill
docs/ architecture, configuration, development, skills
tests/ stdlib-only checks; no pytest, no dependencies
scripts/ build-plugin.sh, bump-version.sh
skills/<name>/SKILL.md 12 skills; frontmatter = name, description (trigger), version
agents/descript-editor.md the autonomous build subagent
references/*.md 4 shared docs the skills read at runtime
docs/ documentation for readers of this repo
dist/ packaged, installable plugin bundle
commands/, skills/, and agents/ are runtime instructions Claude reads. docs/ is for humans reading this repo and is not loaded at runtime.
Once you approve a build plan, descript-editor takes over and runs the passes back-to-back — no approval between steps. It narrates one line per completed pass with the delta (✓ cleanup −1:12), verifies every job against get_project before moving on, and duplicates the composition before the first destructive pass because Descript's AI edits are not cleanly undoable.
It stops for exactly three things: a real (non-transient) failure, an ambiguity it cannot safely resolve, or anything destructive that was not in the plan. A transient failure is retried once.
Individual skills run in checkpoint mode — summarize, then wait. The subagent runs them in autonomous mode. Validation and verification happen in both.
Everything is either a connector you enable in the host or the brand profile cached on first run. There is no config file. Full detail in docs/configuration.md.
- Descript connector — required, OAuth, no API key. Nothing is stored in this repo.
- Figma / brand-voice — optional, used only by brand setup.
~/.descript-studio/brand-profile.md— the one piece of state, kept outside the plugin so it survives reinstalls and is never packaged.
Enforced in the skill bodies, not by configuration:
- Never publishes unless you asked, or publishing was in an approved plan. Transcript export is synchronous and safe, so it runs freely;
publish_projectdoes not. - Always duplicates before the first destructive pass, and says so.
- Never writes outside
~/.descript-studio/and the Descript project you named. - Stops rather than guessing when the target composition is genuinely ambiguous.
No build step, no dependencies. See docs/development.md.
python3 tests/run_tests.py
claude plugin validate .claude-plugin/plugin.json --strict
claude plugin validate . --strictCONTRIBUTING.md covers conventions; CLAUDE.md has the rules that are easy to get wrong. Please read the description rules before touching a skill — descriptions are the routing layer, and all 21 components share one selection namespace.
Security issues go through private vulnerability reporting, not public issues. See SECURITY.md.
| Doc | For |
|---|---|
| docs/architecture.md | How the orchestrator, subagent, skills, and references fit together; the edit loop; execution modes |
| docs/skills.md | Per-skill reference: triggers, MCP tools, pipeline order, guardrails |
| docs/configuration.md | Connectors, the brand profile, defaults, and what the plugin refuses to do |
| docs/development.md | Local setup, the test suite, releasing, and the component-loading trap |
| CLAUDE.md | Rules for agents modifying this repo |
| CONTRIBUTING.md | Conventions for editing or adding a skill |
| SECURITY.md · CODE_OF_CONDUCT.md | Reporting and conduct |
| CHANGELOG.md | Version history |
These are properties of Descript and of AI editing, not bugs to be fixed here — the skills are designed around them:
- Aspect ratio is fixed per composition. "Export 9:16" cannot resize in place; it creates a new reframed composition.
- AI edits are not cleanly undoable. Treat every pass as committing. Before a destructive pass the plugin duplicates the composition so the original survives, and says so.
- Publishing is async, roughly 30s to several minutes depending on length. Transcript export is instant and always safe.
- Filler and retake removal can over-cut. If the result feels choppy, ask for a tighter region or a more conservative pass.
- The first run has no brand to work from. The profile is built on first use and cached at
~/.descript-studio/brand-profile.md; until then the plugin falls back to documented defaults (white captions with a dark outline, lower third, title-card intro, 4s end card) and tells you which it used. Connect Figma or brand-voice, or answer one short round of questions, to replace them.
Every project here shares one idea: a GTM system should refuse to act on data it cannot verify.
cmyk-print-converter — the other production tool here. Drop an image in chat, get back a file a print vendor can actually use.
MIT © Kate Kruger
Descript Studio is an independent, unofficial plugin. It is not affiliated with, endorsed by, or sponsored by Descript, Inc. "Descript" is a trademark of Descript, Inc., used here only to identify the service this plugin integrates with. See the trademark notice.