ArtLab/Local AI: Music Video Experiments

ArtLab · Source archive

Local AI: Music Video Experiments

Three locally generated music videos test a Z-Image-to-M3 production pipeline, an autonomous Hermes run, and the practical problem of reviewing more media than one person can watch.

Originally published as an Advisory Hour Substack post on 2026-08-24. This first-party copy preserves the original text and public supporting media beside the finished work.

Original post

Local AI: Music Video Experiments image from the original Advisory Hour post.
Local AI: Music Video Experiments image from the original Advisory Hour post.

The pipeline is a Z image to M3 video production. Three videos are on this substack from a recent set of experiments. Starcruiser, Feedback Moon, and an Untitled production that a harness took from concept to video without my involvement. These are not award winning AI content.

You’ll watch and listen to these and likely have the same sense I do: none of them are great. At best, they’re mildly interesting. They are, however, a larger step forward in the music video production process. Last week I shared a series of posts where I couldn’t quite achieve a full generative song. One week or so later I’m able to produce music videos based on a generated song. This is a sneak peak of the future. The future is a lot closer than we realize.

Local AI: Music Video Experiments video 1 from the original Advisory Hour post.

The production process does require a harness. I use the opensource harness known as Hermes to help orchestrate all of the steps required. While I do have a tooling website on my local network (LAN) where I can input in all of the finer point settings and controls and generate tracks and media myself, I find it’s not what I gravitate to doing. What I do instead is ask Hermes to go figure it out, and then I go about my day and check on the progress later that day or the following day. The website is useful for review, and Hermes is more useful for creating.

Local AI: Music Video Experiments video 2 from the original Advisory Hour post.

There are still many things to fix. However weird these videos are (and they are weird and would rightfully wear the badge of AI slop with honor) they represent a step-change in capability for what I’m able to do. A lot has happened this year. Since dabbling in local AI due to the AI competitions I’ve gone from generating a few oddball sound effects, to static images, and now it’s possible to generate feature length content. It’s not great content. It’s 3 AM content. Every now and then though there’s a small spark of brilliance.

Local AI: Music Video Experiments video 3 from the original Advisory Hour post.

If I can just do this at home with Local AI, then it seems likely you’ll be able to do something like this on your phone. One prompt and you’ll get a music video, or ten of them. In fact, I can see the problems emerging already in navigating abundance. Here’s the video clips and cast sheet of a Buck Rogers episode I tried. I didn’t quite have the compute to get a viable version built. It is creating a weird problem at home. I now have so many video and music files. I have a couple hairbrain ideas to solve this, but if you think it’s tough to navigate your photo collection now just wait when you really dive into generative media. The asset piles are so large.

Local AI: Music Video Experiments image from the original Advisory Hour post.
Local AI: Music Video Experiments image from the original Advisory Hour post.

Now that I know I can do music videos I’ve set my eyes on a bigger prize. Can I port “John’s World” or even run a feature length TV show (like that Buck Rogers example above). My first experiment in this suggests that it’s possible. It’s possible to generate just about any game now (Mario Walker, if you missed it). It would make sense that TV content would follow. These local AI experiments are sampling from the near future where every pixel is generated.

Connected work

  • MusicLab: Play the growing collection of local tracks and move through it one swipe at a time.
  • Local AI: Music Generation: Hear the earlier Music3 experiments that preceded the finished video runs.
  • Mixed Media Video Production: Watch the earlier long-form experiment built from original music, generated scenes, and archival film.

Use and attribution

Share the work. Credit the maker.

Unless otherwise noted, ArtLab media is licensed under CC BY-NC 4.0.

Commercial reuse needs separate permission. Suggested attribution: Eric Rhea, Local AI: Music Video Experiments, 2026.