r/StableDiffusion 17h ago

Meme can you belive they almost did it t2v

Enable HLS to view with audio, or disable this notification

0 Upvotes

A 10-second 16:9 cinematic shot.

integrated_multimodal_description: [Shot 1] Live-action, cinematic. Dramatic, slightly over-the-top lighting with strong contrast and a cool blue-red comic-book inspired palette, as if inside a stylized Fortress of Solitude or a theatrical stage version of it.

Nicolas Cage appears as Superman — wearing the classic suit with the red cape and the S shield on his chest. His expression and energy are fully in his signature amped-up, wild, unhinged acting style.

[0-4 seconds] He stares intensely into the camera, eyes wide, and launches into the line with rising volume and manic emphasis: “Can you fucking believe that they almost made a fucking Superman movie… with me!… as Superman!”

[4-7 seconds] He throws his head back and laughs maniacally, the laugh loud, sharp, and completely unrestrained, cape shifting with the movement of his shoulders.

[7-10 seconds] He comes down just enough to shake his head rapidly, still grinning with wild energy, and says: “That would have been absolutely terrible.” He immediately breaks into another short burst of laughter while continuing to shake his head in disbelief.

Camera: medium shot that slowly pushes in as his energy escalates, ending in a tighter framing on his face during the final laugh.

overall_soundscape: Clean dramatic room tone with light reverb, his clear, highly animated dialogue delivered in full Nicolas Cage intensity, and his loud maniacal laughter.

non_diegetic_music: N/A


r/StableDiffusion 7h ago

Meme So relaxing, first NVIDIA DLSS 5 Neural Rendering test ( t2v)

Enable HLS to view with audio, or disable this notification

0 Upvotes

integrated_multimodal_description: [Shot 1] Live-action television sitcom style inside Penny’s warmly lit bedroom in her apartment from The Big Bang Theory. A medium-wide shot frames Deadpool lying comfortably on top of the bed in his recognizable red-and-black masked tactical suit, his head resting against the pillows. Penny, played by Kaley Cuoco, sits beside him on the edge of the bed, wearing casual sleepwear. She looks down at Deadpool with an affectionate, slightly amused smile and gently pats his upper chest in a slow, comforting rhythm. The camera pushes in with small amplitude at slow speed as Penny, a young woman with a soft, warm, clear singing voice (S1), sings gently to him: [English] Soft kitty,
Warm kitty,
Little ball of fur.

Happy kitty,
Sleepy kitty,
Purr Purr Purr

Soft kitty,
Warm kitty,
Little ball of fur.

Happy kitty,
Sleepy kitty,
Purr Purr Purr Deadpool remains completely still and listens contentedly, his masked white eye shapes slowly narrowing as though he is becoming sleepy. Penny continues patting his chest in time with the song. As she finishes the final “Purr Purr Purr,” Deadpool gives a relaxed sigh and snuggles deeper into the pillows while Penny smiles down at him. The camera holds on the tenderly absurd bedside moment.

overall_soundscape: Quiet bedroom room tone continues beneath Penny’s singing, accompanied by subtle bedsheet movement, gentle fabric taps against Deadpool’s suit, and his soft relaxed breathing. A faint studio-audience chuckle follows the sight of Deadpool settling sleepily into the pillows.

non_diegetic_music: N/A

https://github.com/lisitskyaa/ComfyUI-DLSS5-NR?fbclid=IwY2xjawUD0lBwZG9mBWV4dG4DYWVtAjEwAGJyaWQRMVM3cUpUUUhaaGZQNmlLZ25zcnRjBmFwcF9pZBAyMjIwMzkxNzg4MjAwODkyAAEeHHM88fwGf57GgXYuPBn_q2TksCGxbf3DFDJhKpn4ZMAR9toPQAJ0GOpaL8g_aem_ZmFrZWR1bW15MTZieXRlcw


r/StableDiffusion 6h ago

Question - Help Looking for Grok img2img Alternative in Local

Thumbnail
gallery
12 Upvotes

Are there any Local Models that can achieve this level of Natural-ness and Realism, not over texturing and over crisp images ? I've been looking for a while and can't find any Closer to this, these images used Grok img2img for Lighting, skin Texture and Overall phone Shot vibes, the base images generated by Local SDXL/Illustrious For the Semi Realistic look, and i used Grok (The Last Grok model before the update), to improve realism, pure img2img and not even a slightest angle change made by Grok, since the Last Grok update everything turned to crap, Everything looks worse and So AI Plastic


r/StableDiffusion 21h ago

Meme Minimax H3, TenStrip eros Beta 4 checkpoint with PK Parasyte 0.5 str, I really like it.

Enable HLS to view with audio, or disable this notification

0 Upvotes

This was made in ~24 minutes total gen time. 1 shot, no redo.

Native res: https://streamable.com/bjwyii


r/StableDiffusion 4h ago

Resource - Update Infinite AI Twitch Streaming Project (2xB200 480p Minimax FastH3)

Thumbnail
youtube.com
0 Upvotes

far from a perfect setup, but this is a interesting project for sure. would not mind the b200 prices to go down tho. framework i used ish https://github.com/reactor-team/infinite-livestream. 2xb200 + gpt-5.6luna


r/StableDiffusion 8h ago

Meme What a waste

0 Upvotes

Nothing worse than waking up to see an overnight run of a like 20 long workflows in Minimax H3 reference only to see everything perfect except the environment is wrong. I have a reference video that has done so well before. I don't see the issue, my prompt is the same as prior success, it's listed in the subject definitions and retention section, and in the prompt too. Oh wait...I didn't link the video to the list of other assets.. ugh.


r/StableDiffusion 15h ago

Resource - Update [Krea2] I trained a Septum Hoop Nose Ring LoKR (Massive shoutout to Fizgig for Win 11 AMD training!)

Thumbnail
gallery
0 Upvotes

Before I get into the details, I want to address the elephant in the room right up front: yes, this is a bit of self-promo, and yes, the model is currently set to paid BUZZ access on Civitai. Please lower your pitchforks!

If you know me, you know I have a full library of LoRAs on Civitai for previous models (SDXL, PonyXL, IllustriousXL, FLUX, etc.) that I have always released for free. I rarely even used the "Early Access" feature. However, with the increased hardware/compute costs of training for Krea2, I’m temporarily using the Buzz system just to recoup those expenses so I can continue producing.

As soon as I hit that break-even point on a release, I will flip the switch and make that individual model 100% free for everyone forever. I have no intention of permanently paywalling my Krea2 work. (I do also have a Patreon, but I don't have any exclusive models locked behind paywalls there either—it's strictly just an alternative way for people to support my work if they choose to). I know limited access is annoying, so I really appreciate you guys bearing with me!

The Problem:
If you’ve been using Krea2/Krea2 Turbo, you already know it’s an absolutely incredible base model. The flexibility and quality are insanely hyped for a good reason. However, it does have some blind spots, which I suppose is expected, and why there's plenty of LoRAs and Fine Tunes online already. One blind spot I noticed: I could never get it to produce coherent Hoop Septum Nose Rings. For me, it almost always defaulted to poor quality horseshoe septum rings (with the gap and balls), would add multiple extra rings, or threw in extra, unprompted, unwanted facial jewelry. Getting a clean hoop natively was basically a nightmare.

The Solution:
I trained a concept LoKR specifically to force clean, highly customizable hoop septum rings.

  • Trigger: SeptumHoopNoseRing
  • Flexibility: Holds up beautifully across all tested art styles, aspect ratios, and distances. Works for men, women, and Character LoRAs.
  • Customization: Fully supports prompting for materials/colors (gold, silver, black, etc.) and sizes (small and thin, large and thick, etc.). Even though it was trained on simple hoops, you can actually prompt it for spiked or jeweled septum rings and it understands the assignment.
  • Settings: Sweet spot is around 0.45 - 0.75 strength using Euler Simple or ddim with a ddim_uniform scheduler.
  • Pro-tip for my model: If you push it to 1.0 strength, it occasionally tries to crop the top of the head. Just briefly describe the subject's hair and eyes in your prompt and it completely fixes it.

The Setup & A Massive Shoutout to Fizgig
I really need to give a massive shoutout to the Fizgig trainer. If you're on a Windows 11 system, especially with an AMD GPU, this tool is an absolute godsend. It finally allowed me to easily train LoRAs/LoKRs locally on my AMD ROCm 7900 XT without jumping through massive hoops or dual-booting Linux.

For the technical folks, here is the under-the-hood breakdown of my training:

  • Dataset: 100 images, hand-refined VL 1st pass captions.
  • Training: 2000 steps using Fizgig’s ultra-fast preset with Adaptive LR enabled and target MP set to 0.5 at ~768px (the max my PC can currently handle).
  • Base: Trained on the FP8 Krea2 Raw model. Examples generated on the Krea2 Turbo NVP4 model.
  • Rig: Win 11, AMD 7900 XT, 32GB RAM, Ryzen 3900x. Tested locally in ComfyUI using the default NVP4 Krea2 model.

Here is the link to the model: [Civitai Link]

What's Next & Commissions
I'm really itching to roll up my sleeves and train some Krea2 models people may have been wanting but haven't seen yet. I'm currently working on re-training my previous models' datasets for Krea2, and plan on releasing several Krea2 models that aren't monetized. I would love to hear your suggestions! I am also open to custom commissions for private use or expedited release (Note: I will not accept commissions or requests to train on IRL persons).

I'd love to hear your feedback or see what you generate with the septum model. If you have any questions feel free to ask!

(Mod Note: I used the "Resource - Update" flair because I didn't know what else to choose. Also, while this is self-promo, I hope it doesn't violate the rules against "excessive self-promo" as I'm aiming to share the resource and training workflow. If this needs to be removed, I completely understand and apologize for any inconvenience!)

(AI Writing Assistance disclosure: The formatting and writing of this Reddit post and my Civitai model description were refined with AI assistance.)


r/StableDiffusion 12h ago

Question - Help krea 2 internet search

0 Upvotes

hey does anyone know of a nodepack or model that connect krea to internet so it can see latest science skeletons instead of relying on frozen knowledge i know of gen searcher but its like 8b i cant run it with krea besides it got no nodes anyway


r/StableDiffusion 9h ago

Question - Help What can I realistically do in Minimax with a 5080

2 Upvotes

I'm getting a 5080, 32gb system RAM

Can I realistically use minimax h3 for i2v and t2v?

How long will generations take. I dont imagine i want to do high quality resolutions. 480p or 720p would be alright


r/StableDiffusion 18h ago

Question - Help Any tips for generating video where people have different accents?

0 Upvotes

I have been employing various tips & tricks from all over Reddit to get consistency and continuation between clips, and I'm in a pretty good spot - or at least I thought I was, until I wanted to generate a clip where an American person is having a conversation with an English person. Then the voices go haywire, the accents get dropped or switched, and gibberish (another problem I thought I'd solved) returns.

Is this just a shortcoming of the model, or is there a trick to generating scenes like this?


r/StableDiffusion 1h ago

Discussion The underrated alternative to krea2 and ideogram4,guess the model?

Thumbnail
gallery
Upvotes

I like ideogram4 and krea2 a lot ,and I also really like this ONE.

What I personally prefer about it is the sense of depth and vastness that I don't feel as strongly in the other two.

Some notes from my testing:

Ideogram 4 can sometimes make skin and details overly sharp in a way that’s hard to fix naturally in post.

Krea 2 occasionally feels a bit static, like the subject was placed into the background rather than existing in the same space (it also uses Wan VAE which is the main reason I have also tested using the fp32 and realvae of it which solved texture to some extent).

These things can of course be improved with better prompting and Loras l,but tendencies are still there.

Just wanted to share a showcase of this model's capabilities.All in all i enjoy all the current models more the merrier.

They all deserve time,testing and appreciation.

some images are made using Boogu base for more creativity,some images are made using boogu Turbo for way greater prompt adherence


r/StableDiffusion 16h ago

Workflow Included MiniMax H3 Ref2VA neural 3D latent upscaling and gentle refinementworkflow Help me push this further on an RTX 4080 16gb 64gb ram

Enable HLS to view with audio, or disable this notification

0 Upvotes

Title: Help me push this MiniMax H3 Ref2VA workflow further on an RTX 4080 16GB

I’ve been building and testing a MiniMax H3 Ref2VA workflow optimized for my RTX 4080 16GB. I’m attaching the JSON and would appreciate help from anyone experienced with MiniMax H3, PDD acceleration, latent upscaling, memory optimization, or continuous video generation.

workflow Download here

What the workflow currently does

  • Uses the pruned INT8 ConvRot MiniMax H3 Ref2VA model.
  • Uses the Qwen3-VL 32B NVFP4/AWQ text encoder.
  • Generates synchronized video and native audio.
  • Uses SageAttention in Auto mode.
  • Applies MiniMax H3 PDD acceleration at 8 NFE.
  • Runs an initial low-resolution PDD render.
  • Separates the video and audio latents.
  • Enlarges only the video latent using the learned MiniMax H3 3D FP16 latent upscaler.
  • Rejoins the upscaled video latent with the original audio latent.
  • Runs a second PDD refinement pass at 0.125 denoise.
  • Decodes the refined video and original audio into an MP4.
  • Includes easy controls for aspect ratio, base megapixels, final target megapixels, and duration.
  • Automatically converts the requested duration into a valid H3 frame count at 24 FPS.

My current general settings are:

  • Base resolution: approximately 0.40 MP
  • Final neural-upscaled target: approximately 0.80 MP
  • Vertical output: roughly 672 × 1216 after upscaling
  • Stable duration: around 5 seconds/124 frames
  • Current workflow default: 7 seconds
  • Second-pass denoise: 0.125
  • Euler sampler
  • Sigma Shift: video 12/audio 3
  • No EasyCache, TeaCache, BlockCache, Spectrum, or additional turbo LoRA stacked on top of PDD

The 0.125 refinement pass only performs about two sampler evaluations in my current setup. At 672 × 1216 and 124 frames, that refinement portion takes roughly 73 seconds.

My current limitations

My practical ceiling appears to be around 7–8 seconds. Going longer causes both my 16GB VRAM and system RAM usage to reach their limits. Five-second clips are currently much more reliable.

I can raise the final target toward 0.90–0.98 MP, but the higher resolution and longer duration quickly increase memory usage. The learned 3D upscaler improves the overall spatial resolution, but it does not reduce the memory required by the high-resolution refinement pass.

My biggest quality issue is facial fidelity. Eyes, eyelashes, skin texture, and other small facial details can still look soft or less refined than they did in my earlier, simpler workflow. Increasing the second-pass denoise too much begins repainting the face, changing the identity, or altering the composition.

My eventual goal is reliable continuous generation. I want to generate several five-second clips by using the final frame of one clip as the starting frame of the next, while still using the original character reference to prevent identity drift.

What I need help with

  1. Is there a better memory-management method for this pipeline that would let me exceed eight seconds on a 16GB RTX 4080 without a major quality loss?
  2. Would model offloading, block swapping, tiled VAE decoding, sequential processing, or another compatible technique reduce peak VRAM and system RAM usage?
  3. Is the learned 3D latent upscaler positioned correctly, or would another order produce better facial details?
  4. Is a 0.125 PDD refinement pass with only about two evaluations doing enough to justify its memory cost?
  5. Is there a better pass-two scheduler, denoise level, or refinement strategy that can improve eyes and skin without repainting the identity?
  6. What is the best way to condition continuation clips using both the previous clip’s final frame and the original reference image?
  7. Are there any H3-compatible face-detail or latent-refinement methods that work temporally and do not cause flickering?
  8. Would decoding/upscaling in smaller temporal chunks help, or would that introduce visible seams and motion inconsistencies?

I’m trying to preserve motion quality, character identity, native audio, and facial fidelity—not simply lower the resolution until it fits.

Hardware: NVIDIA RTX 4080 16GB on Windows 11 using ComfyUI.

Required models are listed inside the workflow notes. I’m attaching the workflow JSON. Any specific node changes, corrected routing, memory settings, or test recommendations would be greatly appreciated.


r/StableDiffusion 9h ago

Question - Help Super realistic human with H3???

0 Upvotes

Wondering has anyone been able to generate super realistic human with minimax H3? I have been trying a lot but the best I got still looks quite AI...

I have seen lots of videos online with super super real human face, the result I got is quite far away from that. So I'm wondering is it limited to Seedance 2.5? Or is there any secret prompt I'm not aware of?

Below is what I mean by super real face I saw online:


r/StableDiffusion 2h ago

Workflow Included Super nothing!

Enable HLS to view with audio, or disable this notification

7 Upvotes

Made with Minimax H3


r/StableDiffusion 2h ago

Animation - Video Hope my humble work would inspire the low vram folks!

Thumbnail
youtube.com
4 Upvotes

An AI-assisted webcomic creator here. I'm among the vram and ram-poor folks, with my humble RTX 3060 12 GB vram and a mere 16 GB ram. Since the beginning of time, I've convinced myself that comic is my focus, and so what I have is enough. I don't want to pay any opportunistic video gen platforms out there. Don't want to rent GPU and trouble myself with transferring assets and models from storage to storage. Aside from light experimentation, I had thought I'd stay away from video gen for a very long while.

That is, until the arrival of Minimax H3... And just two weeks after setting it up (ComfyUI, default ref2va and fl2va workflows), I was able to edit together an animated trailer for my webcomic on my own machine, *entirely local*! Granted, in terms of generation quality there's a lot to be desired, as any resolution beyond 0.4 mp is too slow for me to comfortably iterate on. But still, oh such *feeling* when the world I built suddenly came alive for the first time, and on my own machine, too!

Feel free to ask me anything. Happy to share.


r/StableDiffusion 15h ago

Question - Help What Ai to use to make pictures of an old game look like tripple A games and how to use SD for it

0 Upvotes

I want to make some high-quality images for backgrounds and some uv maps for my favorite PS2 game s.l.a.i. (steel lancer arena international) is stable diffusion the ai i should use for that, and is there something i should know as a noob that has never used it? I do have a 5090 i could use to make images. i just really want to see more content of my favorite childhood game, and AI seems like the only feasible way to get more, so I'd appreciate some advice. I'd like it to make upscaled pictures of stuff like this https://phantom-crash-archive.pages.dev/#/workbench or images i take from the game.

Basically, high-quality renders of a scene using the game objects but reinvented as higher quality.

Sorry if this is an annoying noob question.


r/StableDiffusion 5h ago

Question - Help Which image generator is used for these images?

Thumbnail
gallery
0 Upvotes

Any idea? Is this midjourney?


r/StableDiffusion 6h ago

Question - Help How to improve quality when using H3 with Turbo? (RTX 3000 6GB VRAM)

0 Upvotes

Video link: https://streamable.com/irnjvm

Check video above.

Running MiniMax H3 Turbo on a spare laptop with Quadro RTX 3000 6GB VRAM and 64GB RAM.

352×608, 8 steps, Euler + Beta, Turbo LoRA @ 1.0. Takes ~550 sec for a 5-sec video.

I know the GPU is very limited 😅 Quality isn’t great, but it runs without OOM, so I’m wondering if I can push it further.

Any tips on what to do for better quality? Except for buying a new GPU.


r/StableDiffusion 20h ago

Workflow Included Up at atom! - Behind the scenes of the new Radioactive Man movie

Enable HLS to view with audio, or disable this notification

10 Upvotes

Minimax H3 with turbo lora (default Comfyui template workflow)


r/StableDiffusion 10h ago

Animation - Video Who would win?

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 22h ago

Animation - Video the bird-king (my first fully local AI short film) TW: self-harm.

Enable HLS to view with audio, or disable this notification

21 Upvotes

Minimax H3 baby! It's not perfect and I would love to get your feedback and maybe some tips on how to get rid of plasticky skin.


r/StableDiffusion 9h ago

Animation - Video ALICE MEETS THE RABBIT : REMADE IN MINIMAX H3

Enable HLS to view with audio, or disable this notification

23 Upvotes

About 5 months ago I made clips for a project in LTX 2.3 and remade one of them here in Minimax H3. What a difference a few months makes! Music was created in Suno. I still have to redo some parts with consistency problems but that's enough for today.

The original LTX2.3 version for comparison is here : https://youtu.be/R5tfLKvnJDY


r/StableDiffusion 22h ago

Workflow Included Transformers: Starscream Test #2 - Prompt Below

Enable HLS to view with audio, or disable this notification

3 Upvotes

The classic Transformers series is a blind spot for MiniMax H3, so here’s how I handled this:

System specs: 4070 Ti Super, 16 gb vram, 64 gb ram

Ref2va standard workflow using Fl2va standard model, no loras or speedups.

<Picture 1> Character Sheet

<Audio 1> Vocal Reference

<Video 1> 5 sec Video Reference

Google Gemini to help write prompt.

Added music track in post.

PROMPT:

subject_definitions:

<Subject 1> is the figure in <Picture 1>, featuring a robotic grey face with sharp angular features, glowing red optical visor eyes, a dark grey blocky helmet with side intake vents, and red, white, and blue cybernetic body armor with an orange cockpit chest-plate, blue upper arms and boots, white forearms and thighs, red waist and wing housings, and a purple Decepticon insignia on the wing. Only his robotic design and colors are taken from <Picture 1>; its background, grid lines, and lighting are not carried into the target video. <Audio 1> is the vocal reference for <Subject 1>. <Video 1> is the movement reference; use it as a guide without copying it exactly.

summary:

[reference generation] The target video is a 10-second 2D animated sequence styled after the 1984 Hasbro series The Transformers, featuring <Subject 1> delivering a smug, cutting remark from a metallic Cybertronian battlefield.

retention_analysis:

<Subject 1> (appears in [Shot 1]): fully_preserved - his grey robotic face, glowing red optics, dark grey helmet with side vents, red-white-and-blue cybernetic body, orange cockpit chest-plate, blue limbs, and red wing housings remain unchanged.

detailed_description:

The target video is a traditional 2D hand-drawn animated sequence featuring bold black ink outlines, flat cel-shading, limited animation, expressive poses, and subtle film grain inspired by the visual language of the 1984 animated television series.

[Shot 1] A medium tracking shot frames <Subject 1> from the waist up on a metallic Cybertronian battle platform. He stands with exaggerated confidence, shifting his weight with a classic 1980s cel-animated bounce. He slowly raises one hand, gesturing dismissively toward the chaos unfolding off-screen. His glowing red optics narrow with smug amusement as his wings twitch subtly.

He delivers in the style of <Audio 1>, with sharp, sarcastic timing: <<[English] "Megatron just muted me on the main comms channel. I'd stage a coup, but watching him bumble this conquest is free comedy.">> When <Subject 1> is not speaking he is silent and his mouth remains closed.

On "stage a coup," he gives a brief, knowing smirk. On "bumble this conquest," he gestures toward the battlefield with theatrical disdain. He finishes with a smug stare directly toward camera, holding the pose for a beat before a sharp hard-cel cut.

overall_soundscape:

Metallic servo whines accompany his movements, mixed with distant mechanical explosions, electronic battle alarms, high-tech hums, and echoing Cybertronian machinery.

non_diegetic_music:

N/A


r/StableDiffusion 11h ago

Question - Help Image gen with 9060xt

1 Upvotes

I'm planning on getting a 9060xt 16gb for image generation. I've previously used Automatic1111/Forge with nVidia. I'm wondering if I can reproduce my workflow easily using a 9060xt. I don't mind moving to another application if A1111 is not compatible with AMD but I'd like to know: (a) is AMD compatible with most checkpoints and loras from sites like CivitAI and (b) how fast is generation with AMD compared to nvidia, as in, the 9060xt is comparable to a 5060ti in terms of gaming but can it generate images as quickly?


r/StableDiffusion 1h ago

Question - Help My First AI PC

Upvotes

Hi everyone, I have a question: I'm thinking of buying my first PC solely for AI. What minimum components do you recommend for running Stable Diffusion with Illustrious models? I've been using Free Google Colab to create images in Automatic1111 so I was thinking of buying a PC with similar specifications. What do you recommend?