Top 5 in AI

Guides

How the 'Uptown Funk' Tech-CEO AI Video Was Made (MiniMax H3 on Two DGX Sparks — the First Swap in This Series With a Published Workflow)

By the Top5Apps editorial team · Published October 7, 2026 · Updated October 7, 2026 · 9 min read

Share

Short answer: cut the original video into five-second beats, swap the people in the first frame of each beat with character cards, then feed each beat to MiniMax H3's reference-video mode with that swapped frame as the start — and stitch. That's not our reconstruction; it's what the maker said. On October 6, Gary Lau posted a 2:23 cut of 'Uptown Funk' in which Elon Musk takes Bruno Mars' part and Sam Altman, Jensen Huang, Mark Zuckerberg, Sundar Pichai, Lisa Su, and a man who looks like Dario Amodei make up the band — 'The guys who made RAM so expensive are singing their hearts out… Made on two DGX Sparks using MiniMax H3.' It's at 74,000 views on X, 44,000 more on Bilibili, where he posted it fourteen hours earlier with an explicit AI disclaimer. Three things make it the most instructive clip in this series: it's a full song rather than a 30-second beat; it was generated on desktop hardware with an open-weights model rather than cloud credits; and the maker published the workflow two days before, for a different video, in a seven-minute tutorial. Here's what checks out, what doesn't, and the one license line US readers need before copying it.

AI-generated frame: men resembling Sam Altman in a dark suit, Jensen Huang in a black leather jacket, Elon Musk in a navy suit singing at center, Dario Amodei in glasses and a navy blazer, and Mark Zuckerberg in a grey tee, dancing in a line on a street set with dark storefronts behind them
The group walk, with Musk in Bruno Mars' slot and the others as the band. From Gary Lau's clip, shown editorially; the people depicted had no part in it.
The source, from MarkRonsonVEVO (November 2014; 4:30; 5.91 billion views). Directed by Bruno Mars and Cameron Duddy on the 20th Century Fox 'New York Street' backlot.

What the clip is, and who's in it

The original is a 4:30 street musical — Ronson, Mars, and the Hooligans walking and dancing down a backlot street in late-'70s and '80s clothes, a shoe shine, a limousine, a salon where Ronson and Mars get perm curlers. Lau's cut is 2:23 with the sections reordered, bilingual lyric subtitles burned in, and his watermark top-left. We pulled frames: Musk is the lead in every Mars beat, including the curlers and the limo sprint; Altman (far left, dark suit), Huang (leather jacket), Zuckerberg (grey tee), and a curly-haired man in round glasses who strongly resembles Amodei form the line; Pichai appears from the fourteenth second and ends up in the salon on his phone under a dryer; Lisa Su reads a magazine in the other chair. A young man in a bright blue suit drinking from a flask in one beat is unidentified — he resembles DeepSeek's Liang Wenfeng, the subject of Lau's previous video, but we couldn't confirm it. Ronson himself, the shoe-shine man, and several bystanders are untouched original footage. That mix — swap the principals, leave the extras — is a deliberate economy, and it's in his tutorial.

The joke it's riding

'The guys who made RAM so expensive' is a reply to a reply. On October 3 a Bilibili creator posted Jensen Huang as Michael Jackson on a world tour, with an AI audience of tech CEOs — 2.7 million plays there, 5.65 million views when reposted to X — and on October 5 Musk replied 'This is why RAM is so expensive 🤣🤣' (826,000 views). The line lands because it's true: TrendForce reports conventional DRAM contract prices rose 'approximately 93% to 98%' quarter over quarter in the first quarter of 2026 on AI demand, with another 58–63% expected in the second. The irony Lau may not have intended is that his own hardware is a casualty — NVIDIA raised the DGX Spark from its $3,999 launch price to $4,699 in February, citing 'industry wide memory supply constraints.' The CEOs in the video are, in effect, dancing on a machine their own demand made dearer.

The method, in the maker's words

Lau's October 4 reply, translated from Chinese, is the whole pipeline in one sentence: generate two character images, have GPT split the original MV into shots, replace the people at the start of each shot with the character cards, hand it to MiniMax H3 with the original video as reference, then use the swapped image as the first frame to generate the video. His October 6 tutorial (7:10, Chinese, for the DeepSeek 'APT.' video he made with the same stack) fills in the rest. We read the slides:

  • 1. Character asset card per person: front, side, back, and face close-ups. The same discipline as our Jean Phil guide — the card is attached to every generation.
  • 2. Shot list by Codex. The original is broken into shots, and long shots are 'cut by action, not by fixed seconds' — his examples run about five seconds each. That's why a 2:23 cut is roughly twenty to thirty generations.
  • 3. Swap the first frame. For each shot, an image model replaces the people in the opening frame with the character cards. Extras can stay.
  • 4. H3 reference-video node in ComfyUI. Inputs: the swapped first frame, the original clip as the motion/camera reference, the character cards. He credits the ComfyUI workflow to Bilibili creator AIGC特异点 ('MiniMax H3 Singularity').
  • 5. A six-block prompt, written by Codex from a template: characters and reference definitions; overall summary; what to preserve versus what to transfer; detailed shot description; soundscape; music.
  • 6. Two machines. His architecture slide: a Mac running Codex handles assets, the shot list, and job submit/collect; DGX Spark 1 is the 'head,' DGX Spark 2 the 'worker,' both 'participating in the same video's computation,' exchanging model data over their high-speed link with the Mac out of the data path.
  • 7. QA checklist split into technical, identity and costume, action and props, static, dynamic/audio — then stitch. Asked how long the whole thing took: 'from an idea to a video, one day maybe.'

Does the hardware claim hold up?

Mostly yes, with one caveat. MiniMax H3 is genuinely open-weight. MiniMax open-sourced two variants on August 3 — H3-Base FL2VA (text and first/last-frame) and H3-Base Ref2VA (reference-based, the one this needs) — as a 33-billion-parameter model with native stereo audio, 4–15 seconds a clip, 24 fps, 768-pixel short edge locally. Two pieces stayed hosted-only: Context-IR, the input preprocessor, and Regenerate-2K, the upscaler. The HF card's reference limits: 'Images: ≤ 9 images; Videos: ≤ 3 clips; Audio: ≤ 3 clips.' It runs on a DGX Spark. MiniMax publishes no Spark recipe — its reference configs are 8× B200, 4× H200, or two RTX 5090s at 8 minutes 38 seconds for a five-second quickstart — but an NVIDIA developer-forum package from August 12 reports '~6 min for 5s 720p video' on a single Spark with about 165 GB of models, and a published benchmark got a 15-second 1080p clip with audio down to 741 seconds by generating at 960×540 and upscaling locally. Two Sparks link over ConnectX-7 at 200 Gb/s; NVIDIA's page says up to four can be connected, pooling 256 GB across two. So the pieces exist; whether Lau's two-node split did what his slide says, only his logs would show. One skeptic reply argues the output is too clean for 'the 768p version which is the only open weight version out' — but a local upscale, as in that benchmark, explains 1080p without contradicting feasibility. Budget math: 20–30 clips at 6–12 minutes each is one to six hours of generation before retakes, which fits 'one day.'

The license line US readers need

MiniMax H3's open weights are not licensed for use in the United States. The MiniMax H3 Community License defines 'Excluded Territories' as 'the European Union, the United Kingdom, the Republic of Korea and the United States of America,' requires separate authorization above $20 million in yearly revenue, and requires commercial interfaces to display 'MiniMax H3.' Lau is in China, where none of that applies. For a US, UK, EU, or Korean reader, the sanctioned route to the same model is MiniMax's API or a host that carries it — Fuser's Recast mode, Picsart, Scenario, Pollo — at 4–15 seconds a clip with the same nine-image, three-video reference limits, plus the hosted 2K upscaler the open release lacks. The DGX Spark route is real and the price of two Sparks is about what a year of heavy cloud credits costs; it just isn't licensed where most of our readers live.

Doing it yourself, legally

  • Pick a video with repeatable blocking and a small principal cast. 'Uptown Funk' works because it's five people walking toward a camera for four minutes. The method is shot-by-shot; a video with 200 cuts is 200 generations.
  • Character cards first, invented faces preferred. Seven real executives is the joke here and the exposure (below). The pipeline doesn't care whose face is on the card.
  • Cut by action. Five-second beats that start and end on a pose, not on a timecode.
  • Swap the first frame, then let the reference video carry the motion. This is the single technique that distinguishes Lau's method from the one-click swaps: the model gets the identity from the frame and the choreography from the original, so it isn't inventing either.
  • Keep the extras. Untouched background people are free realism.
  • Write the prompt in blocks. Define each reference by name, say what to preserve and what to transfer, describe the shot, describe the sound. H3's own launch example is the pattern: 'Reference the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matching Audio 3.'
  • QA every clip against the card before stitching. Identity drift between beats is what makes long cuts fall apart.

Rights: seven likenesses, a Sony master, and a disclaimer on one platform

The master belongs to Sony Music (RCA/Columbia); the song has eleven credited writers after the Gap Band credit was added in 2015; the video is a Fox-backlot production directed by Mars and Duddy. Sony is the one major still litigating rather than licensing — it and Universal asserted 61,026 works in their amended complaint against Suno in May, and Sony said last week it had asked platforms to remove more than 260,000 AI tracks by the end of September, calling the fight 'an uphill struggle.' Then the faces: at least seven real people performing nothing they performed, in a clip that puts them under salon dryers. US law on that is in motion — the NO FAKES Act cleared the Senate Judiciary Committee unanimously on June 18 and hasn't passed either chamber; Tennessee's ELVIS Act has covered voice and AI likeness since 2024; California's AB 2655 was struck down last year. Public figures in evident parody is the most defensible version of this, and Lau's Bilibili description does it right: AI derivative work, entertainment only, no participation or endorsement by the people shown, copyrights to their owners. His X post carries none of that. Neither Ronson nor Mars nor any of the executives has reacted.

Our read

This is the first clip in the series where the how-to isn't ours — it's the maker's — and the method is better than the one-click swaps because of one move: swap the first frame, then let the original video drive the motion. That's why a hobbyist with 1,300 followers, four months on X, and two desktop boxes produced a full-song cut that holds identity for two and a half minutes while the template crowd ships 30-second loops. The hardware story is real and worth knowing: a 33-billion-parameter open video model with sound runs on a $4,699 desktop, six minutes a clip. The asterisks are real too: the weights aren't licensed in the US, the 2K upscaler isn't open, and the choice of faces hands the clip's future to seven general counsels and a label that removes a quarter-million tracks a quarter. Learn the pipeline; put your own characters on the cards. Our AI video ranking and open-source models ranking cover where H3 sits against the cloud models.

Where these apps rank