Skip to content

feat(mulmoterminal): clip for renderShapeScript (4.18.0), en and ja - #124

Merged
ystknsh merged 8 commits into
mainfrom
feat/mt-look-before-showing-0911
Sep 11, 2026
Merged

ystknsh merged 8 commits into
mainfrom
feat/mt-look-before-showing-0911

Conversation

@ystknsh

@ystknsh ystknsh commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Summary

4.18.0 の renderShapeScript(エージェントが自分の書いた 3D モデルを 4 方向の 1 枚に焼き、その画像を自分で読む)を紹介する短尺クリップを、英語版・日本語版で追加します。

  • 見せているもの: 机を作らせ、続けて椅子を足させる 2 場面。各場面で、エージェントが焼いたシートを全画面で出す
  • 主役はシート: アプリの 3D ビューア(Canvas)は 4.18.0 より前からある機能なので、payoff はシート側に置き、Canvas は結果として短く添える
  • シートの出所: アプリはこの PNG を表示しないため、撮影中に実際に生成された実物を全画面で挿入し、直前の端末行が名指ししたファイル名とキャプションのファイル名を一致させている
  • 撮影: Codex(gpt-6-astra)のセル。「間違えて直す」ではなく「足す」構成にしたのは、修正型がエージェントの失敗に依存し、モデルが強いほど成立しなくなるため

完成した動画

英語版:

mt-look-before-showing-en.mp4

日本語版:

mt-look-before-showing-ja.mp4

Items to Confirm / Review

  • シートの全画面挿入: アプリが表示しない画像を動画に入れています。端末のファイル名との一致で出所を示していますが、「アプリの画面」と誤解されないかは見ていただきたい点です
  • 英語ビート 2 の余裕が 0.53 秒(クリップ 10.53s / 音声 10.00s)。TTS を作り直したときに伸びると check-beat-fit.py に落ちます。落ちたら chair.mp4 のシート区間を伸ばしてください(ナレーションは短くしない)
  • 尺が 30 秒で、機能クリップの標準(15〜25 秒)を超えています。「何が新しいか」の説明に尺を使う判断です
  • 自己修正ループには触れていません — ツールの指示文は「見て直して焼き直せ」ですが、今回の撮影では一発で正しく書かれ、修正が起きていないためです

User Prompt

  • 4.18.0 が出たので動画を撮りたい
  • 撮影は Codex の astra で実演するのが良い
  • 作った 3D モデルを拡大した方が良い。3 ペインのまま 3D モデルを拡大する意味で。それでも足りなければ映像自体で寄る
  • 「PNG としてこんな風に渡すよ、だからこのように追加も簡単」という流れをイメージしていた
  • 冒頭にモデルが表示されていない
  • ビート間で zoom がリセットされて、また寄るのはおかしい
  • 3D モデルに寄った瞬間に最終ビートへ行ってしまう。2 秒止まってから行けないか
  • 椅子の場面へ繋ぐ一句がないと混乱する
  • TTS は 2.5 pro になっているか

制作メモ

  • 撮影 rig は mulmoterminal-video skill 同梱の record-shapescript.mjs(この PR とは別リポジトリ)
  • 素材テイクは撮影機の footage/2026-09-11/shapescript-g/
  • ナレーションは Codex レビューを 4 周(構成変更のたびに 1 周 + 確認)通しています。主な修正は「必ず誤りを検出する」と読める表現の削除、対話的 3D 操作と誤読される表現の明示化、日本語の尺超過の圧縮
  • 出荷判定(別担当によるフレーム検品)で 5 項目すべて PASS。その後に構成を変更したため、最終版での再判定は未実施です

🤖 Generated with Claude Code

https://claude.ai/code/session_01UxmPTsHfRprS89F9NF4jGH

ystknsh and others added 8 commits September 11, 2026 01:27
The agent writes a desk, renders it, reads its own picture, then is asked
for a chair — which it cannot place without looking at the desk. That is
the whole feature, and it needs no mistake to happen on camera: an
additive second step depends on seeing the first, whatever the model.

The model is enlarged twice: once inside the Canvas with the app's own
zoom, keeping the three panes, and once by the camera at the end of each
scene, because at full width it was seven percent of the frame.

Shot on codex/astra. The ship judgment passed all five gates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UxmPTsHfRprS89F9NF4jGH
The pushed-in frames kept a twenty-pixel sliver of the cell beside the
one being watched, with a status chip sliced down the middle — it read as
a crop that missed rather than a chosen frame. The left edge now lands
past the border.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UxmPTsHfRprS89F9NF4jGH
…already had

The first cut resolved both beats on the Canvas — the 3D viewer, which
shipped long before 4.18.0. The new thing is the PNG the agent renders
for itself and reads, and that appeared only as two lines of tool output.
The author said as much: the clip was hard to follow because the part
being demonstrated was the part already there.

Each beat now cuts from the terminal line naming the file to that file,
full frame, captioned with what it is and where it came from. The
filename on screen and the filename in the caption are the same string,
because it is the same file. The Canvas closes the second beat as the
result rather than the point.

Narration rewritten to match; Codex reviewed it twice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UxmPTsHfRprS89F9NF4jGH
The clip began with the request and the model being written, so nothing
model-shaped was on screen until second eight — including the frame X
shows in the timeline. The author caught it.

A poster beat now opens on the sheet the agent read, and the two working
beats follow it. The order in both is the real one: the agent sees the
model, then you do. The Canvas closes each of them as the result.

Narration rewritten for three beats and reviewed by Codex twice; the
second pass caught that beat one's line ran longer than its own clip.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UxmPTsHfRprS89F9NF4jGH
…ropped terminal

The push-in cut words off both ends of the lines it was meant to make
readable, the top of the panel was empty, and nothing model-shaped was on
screen until second eleven. The author caught it in the rendered file.

The terminal section is now 2.5s at full frame — the request, the call,
the line naming the file — and the seconds it gave up went to the sheet.
Eight and a half of the beat's ten seconds now have a model in them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UxmPTsHfRprS89F9NF4jGH
…g between beats

Two things the author saw in the render. The camera pushed in during each
working beat, so every beat boundary threw the frame back out and crawled
in again — a reset the viewer has no reason for. And the chair arrived at
its full size on the last frame before the cut, so the thing the clip is
building towards was gone the moment it existed.

Every cut is full frame now; the app's own zoom is what makes the model
big, which is the zoom the viewer can see a reason for. A silent beat
holds the finished desk and chair for two and a half seconds before the
end card.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UxmPTsHfRprS89F9NF4jGH
The Canvas updated to the desk-and-chair while the terminal was still
being held, so the finished scene flashed on screen a beat before the
sheet that was supposed to reveal it. The hold now ends at 31.4s of the
take — the line naming the file is up, the Canvas still has the desk
alone — and the chair arrives where it was meant to.

The third beat also began mid-thought; it now opens by asking for the
chair, which is the sentence a viewer needs to follow the turn.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UxmPTsHfRprS89F9NF4jGH
…on asks for

Both decks were copied from a clip that predates the 2026-09-06 decision,
so neither named a model and both fell back to the flash TTS. The author
noticed before they shipped.

Worth knowing while working nearby: several existing clips carry the model
on the _ja deck and not on the en one, so they are still half flash.

The English second beat now clears its clip by 0.53s, which is barely over
the gate — if a later re-render of the TTS comes back longer, lengthen the
sheet in chair.mp4 rather than cutting the line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UxmPTsHfRprS89F9NF4jGH
@ystknsh ystknsh changed the title Feat/mt look before showing 0911 feat(mulmoterminal): clip for renderShapeScript (4.18.0), en and ja Sep 11, 2026
@ystknsh
ystknsh merged commit 72f824a into main Sep 11, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant