r/StableDiffusion 10h ago

Resource - Update VH5 - MiniMax H3 Lora

Enable HLS to view with audio, or disable this notification

311 Upvotes

A style LoRA that makes H3 footage look like it was recorded off 1980s broadcast television onto a VHS tape that has seen better days, soft smeared detail, chroma bleed, tracking noise, head-switching bands at the frame edge, and (because H3 trains audio jointly) the matching muffled mono sound, tape hiss and warble.

https://huggingface.co/KennethFal/vh5tape-vhs-lora-minimax-h3


r/StableDiffusion 3h ago

News MiniMax H3 acceleration arena/leaderbord: 15+ H3 LoRAs, fine-tunes, Max

Thumbnail
huggingface.co
167 Upvotes

Hey folks, I've built an so we can have a proper leaderboard on 15+ different LoRAs, fine-tunes and acceleration technique. Baseline is included for anchoring, and M3 Max is also included given the promise to open source

There are there being compared: H3 baseline, FastH3 family, H3 Acc family, Lightx2v family, Larryvrh family, JoyFox family, RAVEN, FlashGen, TuTu, SilverOxides merges, Plaguekind merges and Fal's H3 Max


r/StableDiffusion 7h ago

Meme My Name Is Giovanni Giorgio

Enable HLS to view with audio, or disable this notification

147 Upvotes

Created with Minimax H3 ref2v using the SEED HUNTER Workflow.


r/StableDiffusion 7h ago

Workflow Included Use H3 To Replace Characters

Enable HLS to view with audio, or disable this notification

82 Upvotes

These characters are very different, so i thought it was a good demo to show. Also, the prompt was an 'omni' prompt which didn't help the model with details on the outfit. Despite that I think it did such a good job wanted to share.

With more details in the prompt related to appearance, acting and dialogue I think you could prob get near flawless changes.

This was don with the FL2VA model, NOT the ref version of H3. It may even be better with the ref but i have been using fl2va mostly because i think the quality is better, but that's subjective.

This concept was inspired by this post originally: https://civitai.red/models/2855941/minimax-h3-character-replacement

I changed the SAM3 use, so technically you could do a multi replacement with some changes. I also added noise to the inverted image, as well as upgraded the prompt to work with FL2VA.

Workflow used to create this video is HERE.

This model seriously continues to amaze me. bravo minimax team, bravo.

Notice in the prompt that the dialogue is the only non 'omni' part at the end, it worked fine not included in the body, since the ref audio is there driving it. again, moving from this general prompt to something more specific i think would give even better results.

PROMPT:
How the reference video and pictures align with the target video — the target video is an edited version of <Video 1>, replacing the silhouette with <Subject 1>.

summary:

[video editing] The target video replaces the silhouette in <Video 1> with <Subject 1>, who performs the exact same motion, dialogue, positions and facial expressions of the silhouette while maintaining the original camera work, environment, and lighting of <Video 1>.

subject_definitions:

<Subject 1> is the person in <Picture 1> and <Picture 2>; <Picture 1> supplies facial features and close-up details, while <Picture 2> provides 3-panel image of front mid shot, profile mid shot, and front full body view, identity follows these reference assets, only appearance is retained.

<Subject 2> is the environment and setting established in <Video 1>. The scene follows this layout, materials, and light; camera position and framing.

<Subject 3> is the silhouette in <Video 1> which provides the motion sequence to be copied.

integrated_multimodal_description:

Video editing, the target video is in a live-action cinematic style with the interior lighting and background and environment atmosphere established in <Video 1> with the likness of <Subject 1> inserted.

[Shot 1] The shot opens with <Subject 1> seamlessly replacing the silhouette <Subject 3> in <Video 1>, the outfit of and clothing of <Subject 1> exactly from reference, performing the exact same motion, dialogue and sounds, positions and facial expressions of silhouette. From the very first frame, <Subject 1> occupies the spatial coordinates of the silhouette replacing with their likness, initiating the same motion onset from rest. <Subject 1> mirrors the silhouette's weight shifts and momentum, body moving in perfect synchronization with the rhythm and pacing of the original footage but replaced with the likness of <Subject 1>. As they navigates the space, <Subject 1> mimics every nuanced gesture—the way the silhouette's head tilts, arm movement, and the micro-movements of facial muscles. The face, defined by <Picture 1>, conveys the same emotional depth as the silhouette, while their full body, as seen in <Picture 2>, provides the physical presence outfit an appearance. The camera follows the exact movement, angle, and cutting rhythm of <Video 1>, maintaining a consistent focal length and distance from the subject at all times. The light from <Subject 2> interacts realistically with <Subject 1>'s skin and clothing, casting shadows that align with the movements of the original scene. The transition is perfect; the result is a fully realized <Subject 1> instead of a silhouette, but the soul of the performance—the timing, the pauses, and the dynamic energy—remains identical to <Video 1>. The movement progresses with a palpable sense of weight as <Subject 1> shifts their center of gravity, with clothes rippling in response to movements. The camera maintains exact framing and cuts as <Video 1>. The scene concludes as <Subject 1> reaches the final position of the silhouette, body settling into a pose that mirrors the original's final frame exactly, with face held in the same expression. <Subject 1> hair, accessories, wardrobe, lighting, and room layout remain unchanged and perfectly replace silhouette throughout.

overall_soundscape:

A low room tone establishes beneath the scene, mirroring the background audio environment of <Video 1>.

<Subject 1> says <d>[English] Can you, can you spare change.</d>.

non_diegetic_music: N/A


r/StableDiffusion 13h ago

Resource - Update MATLOWAI/minimax-h3-fused-turbo-int8-convrot · Hugging Face

Thumbnail
huggingface.co
170 Upvotes

This Minimax H3 all in one checkpoint is quite good.

It merges text, image, and reference to video, as well as 4-step turbo generation into a single model.

No need to switch between models for ref2v, no need to load turbo loras.


r/StableDiffusion 7h ago

News Infinite streaming Slop TV

44 Upvotes

Congratulations everyone! We've done it. Our civilization has reached peak diffusion. It's time to pack up and go home.
https://www.youtube.com/watch?v=EQ2RexjIEFE


r/StableDiffusion 6h ago

Animation - Video HE-MART PSA - MiniMax H3

Enable HLS to view with audio, or disable this notification

29 Upvotes

r/StableDiffusion 17h ago

News A new AI step: Immersive worlds with Minimax H3

Enable HLS to view with audio, or disable this notification

196 Upvotes

The next major interface for artificial intelligence may not be a chatbot, an image, or even a video. It may be a world. 
I designed an H3 Minimax Immersive video workflow for ComfyUI and I want to share it with the Open Community so you can now explore this new field.  

This is an early implementation of that idea using MiniMax H3, using a specialized equirectangular generation ComfyUI workflow with AI 360°prompting to achieve an interactive viewing concept that allows the viewer to control the viewport through the generated environment on mobile and desktop with continuous looping on Youtube and Facebook. 

Watch the immersive demonstration on YouTube. 

The test video is 9 seconds long with a time-reverse layer to get 18 seconds of 360-loop, it was generated using a single 360 prompt

Read my full article with technical data and download the workflow:
https://huggingface.co/blog/zuanfilm/blog

the workflow supports text2-360 and FL2-360, for H3 Minimax 360 prompting I wrote a public custom gpt and added 37000 tokens of 360 filmmaking reasoning 

The result is far from perfect, I generated the clip on my laptop with an Nvidia RTX 3080 Ti 16GB VRAM, so the resolution is very limited and the current generation still shows visible seams on some moments of the video and other inconsistencies but those imperfections may be less important than what the experiment demonstrates. 

Until now a Minimax H3 video was something the viewer has to watch from the camera angle position chosen by the creator, now the viewer can now choose where to look using an immersive UI, that changes the relationship between a person and generative AI media; panoramic video exposes the full spherical observation domain in a single coordinate frame

The generated sequence can be presented as an immersive environment in which the viewer controls the viewing direction. On a phone, the viewer can interact with the scene; on a desktop, the camera can be moved manually. The sequence can also be looped forward and backward so that the environment continues rather than behaving like a single linear cinematic shot. The result is not yet a fully reconstructed 3D universe like a gaussian splatting. It is a time-varying immersive/equirectangular visual environment that can be explored interactively.

The 2:1 rule: the shape of the immersive world

A practical requirement of the equirectangular representation is its 2:1 aspect ratio. For a full spherical panorama: WH=2\frac{W}{H}=2 where WW is the panorama width and HH is its height. For example: W=3840,H=1920W=3840,\qquad H=1920 or: W=7680,H=3840.W=7680,\qquad H=3840. This is the format expected by common 360° video workflows and is particularly important when delivering immersive video to platforms such as YouTube and Facebook where the panoramic video must be interpreted as a spherical 360° environment rather than an ordinary flat video. For example, the H3 generation branch in my workflow uses 2112 × 1056 so the immersive representation and final delivery pipeline preserve the equirectangular 360° geometry.

To manipulate or view the image correctly, computers use 3D rotation matrices.

[ 2D Equirectangular Pixel (x, y) ] 
               │
               ▼  (Convert to Spherical Coordinates)
   [ Latitude & Longitude (θ, φ) ] 
               │
               ▼  (Convert to 3D Cartesian Vectors)
      [ 3D Point (X, Y, Z) ] 
               │
               ▼  <─── MULTIPLIED BY: 3D Rotation Matrix (3x3)
  [ Rotated 3D Point (X', Y', Z') ] 
               │
               ▼  (Project back to 2D)
[ New 2D Equirectangular Pixel (x', y') ]
  • 3x3 Rotation Matrices: These are used to "roll, pitch, and yaw" the camera viewpoint inside the 360-degree sphere. If you drag your mouse to look around a 360-degree YouTube video, a 3x3 matrix is constantly multiplying the pixel coordinates to shift your view.
  • Intrinsic Camera Matrices (K Matrix): A 3x3 matrix that defines the camera's properties—like focal length and optical center. This tells the computer how to crop a normal, undistorted flat perspective view out of the distorted equirectangular image.

This creates an entirely different pipeline: Prompt > AI generation > immersive representation > interactive camera > human exploration The prompt no longer has to describe only what should appear in front of a fixed camera. It can describe a world. That is the conceptual leap, if now this generation process is becoming sufficiently fast, coherent and inexpensive, the applications could extend far beyond experimental video:

Video games Instead of developers manually constructing every environment, AI could generate explorable spaces from natural-language descriptions. “Generate an alien ecosystem surrounding the player.” The difficult question would no longer be only how to render the world. It would be: How quickly can AI generate and maintain the world as the player explores it?

VR education Imagine asking an AI to create an immersive historical environment and then entering it. Instead of watching a documentary about ancient Rome, a student could potentially enter an AI-generated reconstruction and look around. The teacher could change the scenario through language: “Show the city before the fire.” That would transform AI from an information interface into an environment for learning.

AR world transformation The implications become even more interesting when the same concept is combined with augmented reality. A physical environment could become the canvas. A user might look at an ordinary street with some glasses and ask: “Transform this into a cyberpunk city.” “Show this neighborhood as it looked 500 years ago.” or “show me that car in blue with a representation of me as driver” The underlying physical world would remain present, but the AI-generated visual layer could continuously reinterpret it.

Interactive Cinema Movies could eventually become less linear. Instead of the director deciding exactly what every audience member sees at every moment, a film could provide a controlled environment in which viewers explore the scene themselves. The director would still control the story, performances, lighting, world design and narrative boundaries—but the audience could control the camera. That would not simply be another format for film. It would be a new relationship between cinema and audience.

AI worlds driven by AI agents AI agents could eventually generate the environments that humans and other AI agents interact with in real time...


r/StableDiffusion 6h ago

Workflow Included Personification:Planets (tarot cards)

Thumbnail
gallery
21 Upvotes

r/StableDiffusion 22h ago

Resource - Update DLSS 5 Visual Enhancer - standalone neural rendering for images and video

Post image
409 Upvotes

Hey everyone - I made a standalone Windows application for applying a DLSS 5 Neural Rendering feature-18 pipeline to images and video:

Original

DLSS 5

https://github.com/Merserk/dlss5-visual-enhancer

Instead of using DLSS only inside a game, this runs images/video through the ReShade/RenoDX neural-rendering path as a general visual enhancement pipeline.

What it does:

  • Image and video enhancement
  • DLAA/native, 1.5x, ~1.724x, 2x and 3x modes
  • Output up to 8K
  • Neural presets + Natural / Cinematic styles
  • Controls for intensity, local tone, structure and skin structure
  • Batch image processing with before/after previews
  • H.264 / HEVC / AV1 / ProRes video output
  • Video temporal input using optical flow with scene-change resets

GPU support:

  • RTX 40 / 50 series - primary target
  • RTX 30 series - slower beta path

The repository contains the application/pipeline source. Required proprietary and third-party runtime binaries are intentionally not redistributed in the repo.

This is an independent community project and is not affiliated with NVIDIA, ReShade or RenoDX.

I’m especially interested in how this behaves on AI-generated images/video vs normal photography/game footage.

Feedback and comparisons welcome.


r/StableDiffusion 10h ago

Comparison Testing My MiniMax-H3 → LTX 2.5 Upscaling Workflow — Results Are Looking Really Good

Enable HLS to view with audio, or disable this notification

36 Upvotes

I've been testing my MiniMax-H3 → LTX 2.5 upscaling workflow, and the results have been really promising so far.

One thing I've noticed is that the better your original MiniMax-H3 generation is, the better the final upscale will be. I'm getting good results even at lower resolutions, but faces still need stronger and more consistent input generations from MiniMax-H3 to maintain character consistency.

On my RTX 3060 12GB, the current upscale times are roughly:

  • 0.6 resolution: ~15 minutes
  • 0.8–1.0 resolution: ~20–30 minutes

It definitely takes some time, but I'm finding the results are worth it.

And of course, if you have a newer, more powerful GPU, you should be able to get even better results in less time, especially when pushing higher resolutions.

I was planning to release the workflow soon, but I want to spend a little more time testing it and seeing how much further I can improve it before sharing it.

So far, though, I'm really happy with how it's looking. 🔥

Would love to hear what you guys think and whether anyone else has been experimenting with MiniMax-H3 + LTX 2.5 upscaling.


r/StableDiffusion 2h ago

Workflow Included Letting image-to-video artifacts compound into an impossible world

Enable HLS to view with audio, or disable this notification

6 Upvotes

Tools used: Gemma4 12b, LTX-2.3, Wan2GP, vibe coded video editor.

I’ve been experimenting with a slightly self-destructive image-to-video workflow where continuity comes from letting the model reinterpret its own mistakes.

I started with an almost completely black image with a few faint stars, then gave Gemma4 12B the track’s beat grid and energy-shift analysis, along with a long description of the overall concept: a monolith, a hallway of impossible geometry, and a progression from restrained movement into increasingly unstable architecture.

Gemma4 wrote all 27 scene prompts beforehand.

For generation I used LTX 2.3 with the audio-reactive LoRA. I also tested LTX 2.5, but for this workflow it became too artifact-heavy too quickly. LTX 2.3 held the scene structure together longer while still producing enough weirdness to evolve in interesting ways.

The process was simple: generate a clip with the correct audio slice, cut it on the beat grid, then take the frame immediately after the cut and use that as the starting image for the next generation.

The fun part was deliberately keeping some “bad” transition frames.

If a flash landed on the frame used for the next clip, the model might reinterpret it as a permanent light source. A lens flare could become a horizon or an entire landscape. A warped piece of geometry that only existed for one frame could become a major architectural feature in the next scene.

So the artifacts compound.

Eventually the video loses any reliable sense of scale or orientation. Surfaces become spaces, structures fold into other structures, and at some points I wanted an Inception-like feeling where you can’t tell which way is up, or whether the camera is traveling deeper into the structure or pulling outward into something much larger.

The audio-reactive LoRA helps hold it all together. Even when the geometry becomes increasingly strange, the environment keeps breathing, unfolding, compressing and reorganizing itself with the growing low end.

What I like most is that the continuity doesn’t really come from visual consistency. It comes from causality.

Every scene inherits some accidental information from the previous one, and the next generation has to decide what that information actually is.

After enough generations, the model is basically building a world out of its own misunderstandings.


r/StableDiffusion 9h ago

Resource - Update Follow‑up : MiniMax H3 Lip-sync - now does any editable change on a reference video (pose transfer, character swaps, multi‑subject mixes)

Enable HLS to view with audio, or disable this notification

26 Upvotes

Quick update on my earlier audio‑lip‑sync demo: the workflow now chains any desirable edit out of an input reference video for endless video ref pose o lip‑sync.

The audio auto-crop chain is now working for reference videos too and i say it again, I know there are already a lot of options out there for doing this - this is one more option, and it’s definitely not perfect.

VRAM usage went over 40 GB on a 1min run of 3-second, 2MP chucks, so reference-video conditioning is pretty heavy on VRAM and yes you need at least 2MP to get good detail and motion transfer.

Using MiniMax H3’s Ref2V. I’m treating the source clip as the “performance master” (motion, timing, camera) and driving identity/appearance from reference images and audio o the other way around.

What I’ve tested so far: Just MinMax H3 no ControlNet, LoRA, or preprocessor needed.

  • Music‑video pose transfer to new scenarios and characters
  • Single character swap (main performer → reference character) into the ref-video.
  • Multi‑subject mixes:ç
  • Main identity swap
  • Main + 2 added characters, acting in sync or desync
  • Main + 1 added character
  • Replace the main character with 2 characters in pose sync
  • Pull a character from the reference video into an image-reference scene + 1–2 new characters

Everything runs through a single MiniMax H3 chain with mixed references (ref-images + ref-video + ref-audio) and structured prompts that separate identity (image), performance (video), and constraints (text). In practice, every combination I’ve tried is manageable with MiniMax H3.

The node takes the reference video or audio, chunks it into smaller pieces, chains them together, and then stitches everything back together at the end. So, it’s one click, but it can take quite a while to generate a full video.

SUBJECT DEFINITIONS

<Subject 1>: the adult woman visible on the LEFT side of <Picture 1>.<Picture 1> is the appearance reference for Subject 1 only. Its shape, proportion, material, colour, logos and surface markings 100% match <Picture 1>, kept legible and correctly oriented throughout the video.

<Subject 2>: the adult man visible on the RIGHT side of <Picture 1>.<Picture 1> is the appearance reference for Subject 2 only. Its shape, proportion, material, colour, logos and surface markings 100% match <Picture 1>, kept legible and correctly oriented throughout the video.

<Subject 3>: the adult woman main character present in <Video 1>.<Video 1> is the appearance, motion, timing and scene reference.

This is a follow-up to a previous post, so the tips, settings, and links are already available there. MiniMax H3 Lip-Sync: Automatic Long-Video Chaining + Speed & VRAM Optimizations

https://github.com/Ltamann/ComfyUI-H3-Motion-Context-Auto-Chain-addon


r/StableDiffusion 20h ago

Resource - Update [Experimental] DLSS 5 ComfyUI custom node

Thumbnail
gallery
206 Upvotes

Hello Everyone,

Would like to present to you my experimental vibe-coded custom node for DLSS 5 support in ComfyUI.

GitHub project: https://github.com/lisitskyaa/ComfyUI-DLSS5-NR

It's early release, just finished my internal testing and it actually works!

Please note there are no any leaked DLLs in the rep, obtain them separately.

First image in every pair is DLSS 5 ON, second - OFF.

P.S. How to extract original images out of Reddit: https://www.reddit.com/r/StableDiffusion/comments/1p9nrpk/getting_prompt_or_comfyui_workflow_from_posted/


r/StableDiffusion 2h ago

Workflow Included Super nothing!

Enable HLS to view with audio, or disable this notification

8 Upvotes

Made with Minimax H3


r/StableDiffusion 9h ago

Animation - Video ALICE MEETS THE RABBIT : REMADE IN MINIMAX H3

Enable HLS to view with audio, or disable this notification

25 Upvotes

About 5 months ago I made clips for a project in LTX 2.3 and remade one of them here in Minimax H3. What a difference a few months makes! Music was created in Suno. I still have to redo some parts with consistency problems but that's enough for today.

The original LTX2.3 version for comparison is here : https://youtu.be/R5tfLKvnJDY


r/StableDiffusion 1h ago

Animation - Video Made A Professional short Animation video using Minimax-h3 (read description)

Thumbnail
youtube.com
Upvotes

Hey guys!

Since quite a few of you liked my previous videos, I decided to start a channel where soon I’ll be sharing tutorials and some of the workflows/tricks I’ve been using.

If you’re interested in learning how I’m making these videos, feel free to subscribe. I’ll be sharing a lot of the stuff I’ve figured out along the way, including:

  • My own workflows — free to download, with the tricks and settings I use
  • Character generation — how I use a Krea 2 character-sheet LoRA that I made to keep characters consistent, and how to get the style you want
  • Environment generation — how I generate environment images and then build scenes from them
  • MiniMax optimization — settings and techniques to make MiniMax faster while preserving quality
  • Video/audio tricks — ways to fix audio issues and continue a scene from the last frame to create longer sequences
  • Consistent voices — I also built my own UI app using BreezeTTS2 for voice cloning and generating consistent voices across an entire story

Everything I’m sharing is based on what I’ve been experimenting with myself, so hopefully it can save some of you a lot of trial and error.

If that sounds useful to you, you’re welcome to check it out!


r/StableDiffusion 6h ago

Question - Help Looking for Grok img2img Alternative in Local

Thumbnail
gallery
13 Upvotes

Are there any Local Models that can achieve this level of Natural-ness and Realism, not over texturing and over crisp images ? I've been looking for a while and can't find any Closer to this, these images used Grok img2img for Lighting, skin Texture and Overall phone Shot vibes, the base images generated by Local SDXL/Illustrious For the Semi Realistic look, and i used Grok (The Last Grok model before the update), to improve realism, pure img2img and not even a slightest angle change made by Grok, since the Last Grok update everything turned to crap, Everything looks worse and So AI Plastic


r/StableDiffusion 6h ago

Animation - Video Inuyasha Love Triangle Solved

Enable HLS to view with audio, or disable this notification

10 Upvotes

A silly idea I had that I hope you guys had a good laugh at. Still love this classic anime!


r/StableDiffusion 11h ago

Resource - Update Krea2 Turbo Distill 4 step LoRA - new checkpoint (chk42K) released (texture and detail now at 8-step teacher parity, prompt-aware training added, NF4 fully retired for full-int8 training, 1440×1440 now a trained resolution)

Thumbnail
gallery
27 Upvotes

Krea 2 Turbo — 4-Step Distillation LoRA (work in progress)

A LoRA for Krea 2 Turbo that reduces the minimum usable step count from 8 to 4.

  • ⚡ Half the steps — 8 → 4, on Turbo's own deployment sigmas.
  • ⏱️ ~1.6× faster end to end — 54.5 s against the 8-step bar's 88.7 s at 1024×1024, and 1.8× on denoise alone.
  • 🎯 Texture at teacher parity — fine-detail energy 1.00× the 8-step teacher's at 1280×1280 and 1.02× at 1440×1440, matched band-for-band across the frequency spectrum, not grain.
  • 🗣️ Prompt-aware training — the critic scores images against their prompts during training, so adherence is pressured directly, not inherited.
  • 📐 12 trained resolutions — multi-aspect from 512×512 up to 1440×1440, each with its published sweep.
  • 🔌 Drop-in — plain LoRA weights for diffusers and ComfyUI. No custom nodes, no patched sampler, no code.

This is an update release, following up from my previous posts where you can find full details:

Initial, Previous: here, here,  and here

Headline for this update: chk00042000 closes the texture gap: total fine-detail energy against the 8-step teacher reaches 1.00× at 1280×1280 and 1.02× at 1440×1440 (1.0 = teacher-like), where chk00026000 measured 0.88× and 0.82×. And the distribution is right, not just the total — split the spectrum into frequency bands and every band individually lands within ~10% of the teacher's (0.9–1.1×), where 26K ran 0.79–0.92, starved in every band. Total at parity and bands at parity means the detail lives in the same frequencies as the teacher's — real structure, not grain piled into one band. (How can it exceed the teacher? Because the teacher isn't ground truth — training also shows the critic real photographs, so the adapter learns detail density from reality, not only from an 8-step model that itself slightly under-renders fine texture. The teacher anchors structure; reality anchors texture. Values just above 1.0 are that pressure paying off.) In fixed-seed renders it matches chk26K's distance to the 8-step images at 1440×1440 outright.

The recipe grew up since 26K, in four ways: a measured dose of real-image texture pressure — what carried detail to parity; a prompt-aware critic that scores images against their own prompts during training, so effect-heavy prompts now get the energy they ask for; NF4 fully retired — the big resolutions used to squeeze into 24 GB by dropping their attention weights to 4-bit, and after re-engineering the training step to fit full int8, those buckets measure 3.96% closer to the teacher (exactly the buckets texture lives in: 1280², 1440×1280, 1440²); and 1440×1440 promoted to a trained bucket with its own sweep column.

One metric paid for the texture leap — the teacher-velocity score sits a step behind 26K's — a deliberate trade already being won back checkpoint by checkpoint (2.93 → 2.90 → 2.85 and falling) while texture holds parity. _latest now points to chk00042000.

The improvement reaches even the out-of-spec 2-step extreme test. I had a separate dedicated post on that here - since the initial post was done on an earlier to 42K checkpoint, I have since re-rendered the whole native-vs-LoRA 2 step strength-2 set on this checkpoint (42K being released now), and the FFT is the diagnostic: the old 2-step had the classic collapse signature — hollow mid-bands (0.52/0.55) plus a fake-grain overshoot at the very top (b6 = 1.05). This checkpoint lifts every structural band (0.64/0.65/0.76/0.80) and settles the top band to 0.82 — more real structure, less noise dressed as detail. Fresh strips: 2-step extreme test. And that's the preview mode (at quick 2 steps, unofficial, untrained for, still useful for previews, and getting better and better with every new checkpoint release).

Which file to download

file use it when
krea2_turbo_4step_rank_64_lora_latest.safetensors normally — always the newest accepted checkpoint
krea2_turbo_4step_rank_64_lora_chk00042000.safetensors pin this exact checkpoint

and, beside them, the same files with a _comfyui suffix for ComfyUI. Earlier checkpoints (chk00004000chk00005000chk00006000chk00010000chk00014000chk00019000chk00026000) are kept in older_checkpoints/, and their resolution sweeps stay in place, so the progression remains visible and comparable.

For the full 42K Checkpoint resolution sweep go here: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA/tree/main/checkpoint_resolution_sweeps/chk42000

This is work in progress and better checkpoints may follow. Training is ongoing, so ..._latest... is a rolling pointer: when a newer checkpoint is accepted, that filename gets the new weights and a new numbered copy appears beside it. Re-download the _latest file and everything keeps working — the ComfyUI workflow references it by that name (it does get updated Note in it so technically it is updated but not functionally). Pin a numbered file instead if you need reproducibility.

How checkpoints get chosen

This is not a "train for longer and ship the newest file" project. More samples do not reliably mean a better adapter — measured here, they can make it worse, and a higher number on its own means nothing.

The loop is train → assess → adapt the recipe → retrain → assess again, and a checkpoint is published only when it is measurably better than the one it would replace, on the same held-out set and the same evaluation, and its full resolution sweep shows no regression. Runs that come out flat or worse are kept as information about the recipe and discarded as releases — several have been.

So the recipe itself changes between runs. Each published checkpoint reflects whatever the previous round taught us: the training precision, the optimiser settings, the teacher used to generate the targets and the data mix have all been revised on evidence rather than assumption.

Timeline of training process

Each checkpoint is the product of several stages with very different costs:

  1. Text-encoder embeddings. Every training prompt is encoded once and cached. This is the fast part — thousands of prompts take minutes.
  2. Teacher shards. For each cached prompt, the unmodified Krea 2 Turbo runs its full 8-step schedule and the whole trajectory is recorded, at every one of the supported resolutions. This is by far the most time-consuming stage — it is the teacher doing real inference, thousands of times, and a batch of several thousand shards is measured in days of GPU time, not hours.
  3. Real-photo crops. Bucket-sized crops are cut at native resolution from quality-gated real photo sources (public high res datasets), VAE-encoded into the training latent space, and captioned per crop for the prompt-aware side of training. Cutting, encoding and captioning a pool refresh is a matter of hours.
  4. Student training. The LoRA trains against the recorded trajectories (progressive distillation), with a latent-space GAN critic running alongside — real crops and teacher finals as its real class, the student's outputs as fake — plus a prompt-aware head that scores images against their prompts. Relative to the shard stage this is quick: each +1,000 checkpoint is a matter of hours, not days. Of course the longer the training the better and more diverse results, so hours do turn into days eventually.

Full details and to download - check my Hugging Face LoRA

HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA


r/StableDiffusion 4h ago

Discussion Thoughts & opinions on Anima - turbo-v1.1

Thumbnail civitai.red
5 Upvotes

So I've been using the new Anima turbo-v1.1 model and I have to say it's pretty good now and then. The thing I like is its unpolished look like it does have a rough default art style in my opinion, but I kind of like that as it looks less too polished. It also has pretty good diversity as well. It also works pretty well with LORAs like the base model, however I haven't tried multiple LORAs together.

What I don't like about it is it can be a little inconsistent regarding prompt adherence and also Anatomy and sometimes it can give it for the subject extra or missing limbs and miss out details/objects in the prompt sometimes. Not often but sometimes. To be fair, the turbo model also has the same issue as well sometimes and is probably due to the low CFG and low steps of turbo distilled version.

It looks like a bit of an improvement to the previous version, but I do hope the anima team works on a bigger and stronger turbo Lora for the base model as it's still much better especially when using other fine-tune anima Checkpoints plus better Lora support.

I'm curious to see what you guys think of it as it is a fairly new release.


r/StableDiffusion 6h ago

Question - Help Minimax H3 Ref key words help

11 Upvotes

I've read through the prompt guide, but I'm still having some trouble understanding when to use which of these

fully_preserved, partially_preserved, attribute_transfer, weak_reference

From what I understand you use these in the retention_analysis block. Let's say I want to fully_preserve the face, hair, and body characteristics from <Picture 1>, but I want to swap the character to wear the clothing from <Picture 2>.

Do I use

<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - and describe the portions of the picture I want to fully_preserve? 
<Picture 2> ([Shot 1] first frame): fully_preserved - and describe the clothing I want to fully preserve?

or do I 

<Subject 1> (appears in [Shot 1], [Shot 3]): partially_preserved - because I want to change the clothing she's wearing?
<Picture 2> ([Shot 1] first frame): partially_preserved? Or attribute transfer?

r/StableDiffusion 2h ago

Animation - Video Hope my humble work would inspire the low vram folks!

Thumbnail
youtube.com
3 Upvotes

An AI-assisted webcomic creator here. I'm among the vram and ram-poor folks, with my humble RTX 3060 12 GB vram and a mere 16 GB ram. Since the beginning of time, I've convinced myself that comic is my focus, and so what I have is enough. I don't want to pay any opportunistic video gen platforms out there. Don't want to rent GPU and trouble myself with transferring assets and models from storage to storage. Aside from light experimentation, I had thought I'd stay away from video gen for a very long while.

That is, until the arrival of Minimax H3... And just two weeks after setting it up (ComfyUI, default ref2va and fl2va workflows), I was able to edit together an animated trailer for my webcomic on my own machine, *entirely local*! Granted, in terms of generation quality there's a lot to be desired, as any resolution beyond 0.4 mp is too slow for me to comfortably iterate on. But still, oh such *feeling* when the world I built suddenly came alive for the first time, and on my own machine, too!

Feel free to ask me anything. Happy to share.


r/StableDiffusion 14h ago

Question - Help What's the fuss with hybrid Minimax H3 models ?

36 Upvotes

I don't understand the trend of hybrid models (ref2va blocks over fl2va)

It's supposed to have the best of both worlds : reference adherence through the refva2 blocks and best quality through fl2va as fl2va is supposed to have somewhat better quality

Well my experience so far, and I hope it's a skill issue to be honest, is that the reference part is much less random and unprecise... and for the quality gain i'm not sure, and anyway it's pointless if the video rarely respect my references or starting pic.

Even using a keyframe guide as the first pic I find often the video only using it at first and immediately switching to something else, or the opposite, following the prompt after inserting a random pic at first. Some stuff like that.
(At least fl2v always respect first and last frame)

Not sure if it's due to accelerating stuff or not, as I've tried some hybrid models with 25 steps as well and it was more or less the same

Am I doing something wrong ? Do some people have the same experience ?

I'm asking that because it wouldn't be the only time there's a buzz on something and we just didn't hear the opposite experiences (for example we have been told a LOT of times spectrum doesn't degrade anything but after playing many times with it, even trying conservative settings, I got rid of it, as it WAS degrading things... mileage can vary)


r/StableDiffusion 20h ago

News Trellis.2 and Pixal3D Are Now Native in ComfyUI

Thumbnail
gallery
97 Upvotes

Both Trellis.2 (Xiang et al., 2025) and Pixal3D (Li et al., 2026) now run natively in ComfyUI. No custom nodes, no compiled CUDA extensions, no PyTorch downgrades, and no non-commercial dependencies.

This is more than a model integration. It ships with a rebuilt 3D pipeline: new Load/Preview/Save 3D nodes, a set of mesh post-processing nodes, and an extended PBR texturing stage that bakes normal and ambient occlusion maps for a complete material set. Everything runs on consumer hardware, and everything is free to use, including commercially.

Why Trellis.2 still matters, ten months later

When Microsoft open-sourced Trellis.2 in December 2025, it immediately became the best open-source model for 3D generative AI. A 4-billion-parameter model built on a compact structured latent representation (O-Voxel). It generates high-fidelity 3D assets from a single image at effective resolutions up to 1536³, handling complex topologies that earlier methods struggled with. It also shipped with a PBR texturing model generating base color, roughness, and metallic maps.

Ten months is an eternity in generative AI, yet Trellis.2 hasn’t just aged well, it has become foundational. Several open-source 3D models released since build directly on it, the most notable being Pixal3D whose implementation uses the Trellis.2 backbone.

The community got there first

As always, the ComfyUI community was quick to bring Trellis.2 into the graph. Within days of the release, custom node packs appeared, the most popular being ComfyUI-TRELLIS2 by Andrea Pozzetti and ComfyUI-Trellis2 by VisualBruno, which together gathered well over a thousand stars. We’re grateful to both authors as they proved the demand and carried the community for months.

Despite their efforts, running Trellis.2 remained a challenge for two reasons.

Installation

The original implementation targets environments built around PyTorch 2.6.0 with CUDA 12.4, which for many users meant downgrading their existing ComfyUI environment. On top of that sit a stack of compiled CUDA extensions (flash-attention, FlexGEMM sparse convolutions, the O-Voxel kernels, CuMesh, nvdiffrast) each of which must match your exact Python, PyTorch, and CUDA combination. The custom node authors did heroic work shipping prebuilt wheels per configuration, but every PyTorch or CUDA update meant a new round of compilation failures, and installs regularly broke. This is now solved with the native integration in ComfyUI. Follow our installation tutorials for Trellis.2 and Pixal3D.

Licensing

Trellis.2’s own code and weights are MIT-licensed, but its original pipeline depends on NVIDIA’s nvdiffrast (for mesh rasterization) and nvdiffrec (for Physically Based Rendering), both distributed under the NVIDIA Source Code License which restricts usage to non-commercial research and evaluation. In practice, a studio couldn’t ship assets from the reference pipeline without stepping into a legal gray zone. These dependencies have been removed from with the native integration.

Then came Pixal3D

In April 2026, Pixal3D from researchers at Tsinghua University and Tencent ARC Lab got accepted at SIGGRAPH 2026. It pushed open-source 3D generation another step forward with its pixel-aligned generation establishing direct pixel-to-3D correspondences. The result is near-reconstruction-level fidelity to the input view, with detailed geometry and the same PBR material set.

Pixal3D is heavily built on Trellis.2 as it uses its backbone and shares its VAEs and DINOv3 image conditioning. This is why integrating it together with Trellis.2 made sense. However Pixal3D generally performs better than Trellis.2 as the generated 3D mesh strictly aligns with the input image.

Model highlights

Trellis.2

  • Single image to 3D asset. A 4-billion-parameter model that generates high-fidelity geometry and materials from one input image.
  • O-Voxel structured latents. A native, compact omni-voxel representation encoding both geometry and appearance, generating assets at effective resolutions up to 1536³.
  • Any topology. Handles open surfaces, non-manifold geometry, and fully-enclosed volumes.
  • PBR materials built in. A dedicated texturing model generates base color, roughness, and metallic maps.

Pixal3D

  • Pixel-aligned generation. Geometry is generated in direct correspondence with the input view. What you see in the image is what you get in 3D!
  • Explicit image back-projection. Multi-scale image features are lifted into a 3D feature volume, delivering near-reconstruction-level fidelity.
  • Cascaded refinement. A staged process progressively refines sparse structure, shape, and texture up to high resolution.
  • Built on Trellis.2. Shares the Trellis.2 backbone, VAEs, and DINOv3 conditioning.

What ships in this integration

The goal was simple: make the best open 3D models run in ComfyUI the way every image or video generation model does. A major thank-you goes to Kijai for the implementation, and to yousef-rafat for the initial draft this work built on. In addition to the native implementation, this has been an opportunity to make 3D generation a first-class citizen in ComfyUI. Here is what shipped:

Pure native implementation

Both Trellis.2 and Pixal3D now run as core ComfyUI nodes. The 3D post-processing that required compiled extensions has been reimplemented from scratch in PyTorch and SciPy. No nvdiffrast, no nvdiffrec, no per-configuration wheels, no PyTorch downgrade. If your ComfyUI runs, these models run on your current PyTorch.

Rebuilt 3D nodes

While these were shipped in an earlier version of ComfyUI, the Load 3D, Preview 3D, and Save 3D nodes have been rebuilt from the ground up to support these models and modern mesh workflows. We’re grateful to Terry Jia for his remarkable work on these nodes. Check out the nodes:

  • Load 3D (Advanced)
  • Preview 3D (Advanced)
  • Save 3D (Advanced)

Native mesh post-processing

Raw generative meshes are rarely production-ready, so this release introduces a new set of post-processing nodes:

  • Remesh Mesh: fixes holes and mesh imperfections.
  • Decimate Mesh: reduces face and vertex count to a target budget.
  • Smooth Mesh Normals: smooths the mesh volume.
  • Fill Holes: fill-in holes resulting from the generation
  • And more: Merge Meshes, Paint Mesh, Render Mesh…

A complete PBR texture set

Trellis.2’s texturing model generates base color, roughness, and metallic maps. Our implementation goes further: a new UV unwrapping node prepares the mesh for texturing, and two additional maps are generated: a normal map and an ambient occlusion map, both baked from the high-poly mesh. Are these textures perfect? No. But they’re free, generated on consumer hardware, and yours to use as you wish.

An honest word on quality

Let’s be direct: the best closed-source 3D generators (Hunyuan 3D, Tripo, Rodin) still produce better results than Trellis.2 and Pixal3D. If you need the highest quality and an API fits your pipeline, those remain strong options (all of them are available through ComfyUI’s partner nodes).

What this integration offers is different: the best open 3D generation available, running locally, at zero cost per asset, with no licensing restrictions on what you make. For iteration, prototyping, stylized work, 3D-to-2D workflows, and anyone who wants full control of their pipeline without spending an afternoon to install.

Getting started

  1. Update ComfyUI to the latest version 0.34.0 or go to Comfy Cloud
  2. Download the workflows below, or find them in the template library.
  3. Follow the note in the workflow to download the models and save them in the correct model directory.
  4. Drop in an image and run.

Download Workflow

Model weights: