Model Training · Source archive

Local AI: Minimax H3 on a Nvidia 3080

Originally published as an Advisory Hour Substack post on 2026-08-10. This first-party site copy preserves the text, localizes public supporting media, and connects the record to its project trail.

Original post

Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.
Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.

Squirreled away on a hand-made rack strung together by wires and blinking lights is a little machine. A machine the size of two stacked toasters that in hindsight was a strategically rare purchase: a liquid cooled 3080. Today’s article is about a new breakthru on this old hardware. The kind of thing I intuited would be plausible-just when? Video and audio generation of reasonable quality.

Running a Minimax H3 on my NVidia 3080 solves a major question mark of capability for my local AI systems that until now was impossible. Could I ever generate video? Doing images and text are great and I’ve written so many articles on what sorts of tricks I’ve found can be done with those. What about images in motion? Would I ever reach a day when video might be at all possible? On a rainy august day, it seemed unlikely-but then everything changed.

I applaud the technical breakthru the team over at Minimax achieved with this model. It just opens up so many great use-cases for me-experiments that I thought were forever gated behind a paywall. Instead, I can generate long duration, high quality video for the cost of electricity.

The First Video

After a few system adjustments came the moment of truth: what would be the very first video? A complex scene? Some scandalous IP involving the cast of friends stranded on Hoth? I needed

What I was after was a test of physics. Could this model handle physics at any reasonable level? If so, then it is a massive unlock. Up until now the video I can generate on my home hardware is so lackluster I’ve not even bothered to post it. The videos blend and blur in ways that make even melted crayons look like a work of art.

Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.
Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.

The very first video was a simple test. Let’s try rolling a ball. This may sound like a ridiculous thing. Surely all video models can do this! You would be mistaken. Not all videos can handle this task. What happens is the ball merges into the scene, melts, or otherwise warbles. This time, however, it worked.

The Second Video

The second video was an expanded test on the idea, with audio. Could I generate a video with audio featuring a red ball? The hardware I have available released nearly six years ago. Generating a video with sound felt like an impossibility. And yet? That, too, works. The easiest way to test audio and a ball is to hear the impact of each bounce.

The contact sheet generated from the video shows each and every bounce. It’s quite interesting to review. Having the entire power of a world simulator on my 3080 is such a delightful surprise.

Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.
Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.

I also generated this chart as a probe on what the ball simulation did. Here you can see each bounce, as well as GPU thermal properties. Heat is a very real problem. I have a system that’s liquid cooled, and it still runs hot.

Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.
Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.

Let’s briefly segway into hardware.

Heat and Hardware

I have Hermes monitoring the job runs. This gives me an easy to run “swiss army knife” if something goes sideways. And in this line of work? Something almost always goes sideways. Note that I have an emergency fail safe to abort a run if the gpu is running too hot. I need to put cycles into optimizing H3 on my hardware, but until I do that the thermal abort is my way of keeping the hardware safe.

Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.
Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.

Video Models Mock Physics Algorithms

A quirk of video models that I feel not enough people have explored is the unexpected way in which it takes the same time to achieve one thing over another thing. For example, it takes the same amount of time to simulate 1 ball dropping as it does 1 million. As someone who has put countless hours into optimizing all sorts of algorithms, it drives me a special kind of insane. What magic is this?

The contact sheet for this video:

Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.
Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.

The video generation is not 4k. I am not capable right now of 4K video. I have ideas on how to possibly pull that off, but I’m saving that as an experiment to write about later on. It’s hard to express how delighted I am by these results.

Up until now, the best video I could hope for was 1.5 second blurs that made no sense. Now I have audio and physics, and a machine that I can put to work. In fact, it opens up new wonders of exploration that are just waiting. If you followed my John’s World series, then you can imagine how excited I am that I can use my local hardware and make this scene happen. A simple trainride, with audio, out of Chicago.

Contact sheet:

Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.
Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.

If you’re someone that enjoys looking at performance charts, then check out the following details about the hardware profile. This is using an adaptive cooling solution driven by software-not hardware. The cooling profile maintains a far more reasonable usage than I ever saw dabbling in web3. If you connect this chart into your favorite AI tool, then you’ll be able to entertain a wide conversation.

Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.
Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.

It is incredible to me that I can generate a trainride out of chicago after dropping a million red balls over times square. And I drove all of this on local hardware? Outstanding. Now what else can I do? I’m excited to find out.

If prompts are your thing, here’s the prompt.

A single continuous cinematic documentary shot from a fixed passenger seat inside a moving intercity train, looking out through one large clean side window as the train rides out of Chicago in late-afternoon summer light. The dark window frame stays rigid and anchored at the far edges of the image while the exterior moves continuously from right to left with physically coherent forward-travel parallax. At the beginning, the recognizable Chicago downtown skyline including the dark Willis Tower recedes behind converging railway tracks. The train passes broad rail yards with parallel steel rails, switches, signal gantries, weathered brick warehouses, steel bridges, utility poles and modest Chicago neighborhoods, gradually transitioning toward lower-density industrial outskirts while the same skyline shrinks naturally in the distance. Keep track geometry straight and connected, buildings stable, and motion speed constant. Subtle realistic reflections slide across the window glass without obscuring the view. No cuts, no jump transitions, no camera reset, no time lapse, no impossible track warping, no readable sign emphasis, no people inside the carriage. Generate continuous synchronized clean stereo audio: steady steel wheel-on-rail rhythm and track clacks matching the train speed, low carriage rumble and vibration, light wind at the window, occasional distant rail squeal, one restrained faraway train horn, and fading Chicago city ambience; no music and no narration.

ChatGPT will create an image with that prompt that looks something like the image below. It’s a nice view and shows the level of detail missing from the lower resolution video.

Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.
Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.

Upscaling the video of the trainride itself is an interesting thought. What if, one might think, you took the lower resolution video then ran literally every frame thru a 4k upscaler system? After all, I did write about upscaling different images for Artlab from time to time.

What I see is the upscaler struggling with missing detail. Absent anything to fix the missing detail, the entire system starts hallucinating pixels. A worthwhile experiment here would be to take each frame and try using img2img transfer in a very nuanced way in an attempt to layer in that missing detail (like a zoom enhance kind of idea).

Here’s an incredible scene from a TV show in the mid to late 1900s called Taxi. I remember watching this on a three dial television. Even crazier is the raw image format. This is using base64 encoded webp image data that I tossed into a file in the root drive, then pointed Hermes at the file and said, “Surprise me.”

Born to Bounce?

This entire article began with a red ball and then we ended with a never-existed before scene from the television comedy show Taxi. ChatGPT suggested today’s post is sponsored by a sticker for those born to bounce, and I agreed. Pickup your sticker for today’s post at Bubble-free stickers - Born to Bounce

Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.
Local AI: Minimax H3 on a Nvidia 3080 image from the original Advisory Hour post.

Connected work