MiniMax-H3 通过 MLX 移植可在 Apple Silicon 上运行
PipeNetwork/minimax-h3-mlx
MiniMax 发布 MiniMax-H3,一个可接受文本、图像、音频和视频并生成最长 15 秒带音频视频片段的通用全模态生成系统。Python 包 PipeNetwork/minimax-h3-mlx 将其移植到 MLX,支持 Apple Silicon 运行。作者在 M5 Max MacBook Pro 上实测,下载约 115 GB 模型文件,视频生成耗时不到 45 分钟。
在 M5 Max 上生成本地视频的 45 分钟耗时,为判断这类模型何时能进入日常工作流提供了具体的参照。
two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video and can use them to generate up to 15 second video clips with audio included.
This Python package ports it to MLX for running on Apple Silicon.
I got it running on my M5 Max MacBook Pro. I cloned the repo and ran the model like this:
# First download the models
uvx --from huggingface_hub hf download MiniMaxAI/MiniMax-H3 \
--include 'FL2VA/*' --exclude 'FL2VA/transformer/*'
uvx --from huggingface_hub hf download pipenetwork/MiniMax-H3-MLX-8bit
# Now run the prompt
uv run --with mlx-vlm \
--with-requirements requirements.txt python scripts/generate.py \
"a rainbow colored skunk leaps over a mossy log in a supermarket" \
-o skunk.mp4 \
-c ~/.cache/huggingface/hub/models--MiniMaxAI--MiniMax-H3/snapshots/fa9c8ab1eaa21c8ae25e7e40b83b2e6002f340af/FL2VA \
-t ~/.cache/huggingface/hub/models--pipenetwork--MiniMax-H3-MLX-8bit/snapshots/3ac52081470b0488921c3ec3ba84a39097bf2361
Here's the video I got for the prompt:
a rainbow colored skunk leaps over a mossy log in a supermarket
视频 · 前往原文观看It downloaded ~115 GB of model files, and the video generation took just under 45 minutes.
The video is impressive, but the audio is weird speech-like garbage, because I didn't provide any prompt guidance as to what the audio should be. The prompting guide (which I didn't read prior to this experiment) has a whole bunch of information on how to get this to work.
Tags: ai, generative-ai, mlx, text-to-video, minimax
来源:Simon Willison 博客 · simonwillison.net