Skip to content

feat(text chat): expose image input via repeatable --image (multimodal M3, multi-image) - #225

Open
shoemoney wants to merge 1 commit into
MiniMax-AI:mainfrom
shoemoney:feat/text-chat-image-flag
Open

feat(text chat): expose image input via repeatable --image (multimodal M3, multi-image)#225
shoemoney wants to merge 1 commit into
MiniMax-AI:mainfrom
shoemoney:feat/text-chat-image-flag

Conversation

@shoemoney

@shoemoney shoemoney commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Closes #224.

What

mmx text chat now takes a repeatable --image <path-or-url>:

# single image
mmx text chat --image ./photo.jpg --message "What breed is this dog?"

# multiple images in one call
mmx text chat --image ./before.png --image ./after.png \
  --message "List every visual difference between these two."

Why

M3 is multimodal and already accepts multiple images through --messages-file, but there was no CLI path to it — you had to hand-write a base64 messages JSON. Worse, the documented OpenAI shape ({"type":"image_url", ...}) is rejected, because text chat posts to the Anthropic-compatible /anthropic/v1/messages endpoint and needs {"type":"image","source":{"type":"base64",...}}. --image emits the right shape for you.

How

  • toImageBlock() in src/utils/image.ts wraps the existing toDataUri() (local paths, http(s) URLs, and pre-made data URIs all work) and returns an Anthropic image block.
  • src/commands/text/chat.ts appends the blocks to the last user message, promoting content from a string to a block array. Text goes first so the model reads the instruction before the pixels. With no --message, the images become the user message.
  • ContentBlock gains the image variant.
  • When --image is present and --model is not, the model resolves to MiniMax-M3 — otherwise a text-only defaultTextModel in config would silently break every image request. An explicit --model still wins.
  • Docs updated: README.md, README_CN.md, and skill/SKILL.md (which also now calls out the OpenAI-vs-Anthropic block-shape gotcha).

Verification

Dry run of the built binary:

$ mmx text chat --image a.png --image b.png --message "diff these" --dry-run --output json
{
  "request": {
    "model": "MiniMax-M3",
    "messages": [ { "role": "user", "content": [
      { "type": "text",  "text": "diff these" },
      { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "..." } },
      { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "..." } }
    ] } ],
    "max_tokens": 4096,
    "stream": false
  }
}

That is byte-for-byte the payload shape I confirmed M3 answers correctly in #224.

Six new tests in test/commands/text/chat.test.ts cover block shape, multi-image, image-without-message, the model override, explicit --model winning, and the missing-file error. bun test 455 pass / 0 fail; bun run typecheck and bun run lint clean.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

MiniMax-M3 is multimodal, but `text chat` had no way to send an image —
users had to hand-write a base64 messages JSON file, and the obvious
OpenAI `image_url` shape is rejected because the CLI posts to the
Anthropic-compatible /messages endpoint.

- `--image <path-or-url>` on `text chat`, repeatable, so multi-image
  compare/diff works in one call
- new `toImageBlock()` in utils/image reuses `toDataUri()` (local paths,
  http(s) URLs, existing data URIs) and emits the Anthropic block shape
  `{ type: 'image', source: { type: 'base64', media_type, data } }`
- images append to the last user message, promoting string content to a
  block array; with no `--message` they become the user message
- images force `MiniMax-M3` when `--model` is unset, so a text-only
  `defaultTextModel` in config can't silently break the request
- docs: README, README_CN, skill/SKILL.md (incl. the OpenAI-vs-Anthropic
  block-shape gotcha)

Closes MiniMax-AI#224
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

text chat: expose image input (enable multimodal M3, incl. multi-image)

1 participant