Skip to content

SimpleTuner

SimpleTuner is a multi-modal diffusion model fine-tuning toolkit focused on simplicity and ease of understanding.

Features

  • Multi-modal training - Image, video, and audio generation models
  • Web UI & API - Train via browser or automate with REST
  • CaptionFlow captioning - Generate captions with local GPUs through the Web UI job queue
  • Worker orchestration - Distribute jobs across GPU machines
  • Enterprise-ready - LDAP/OIDC SSO, RBAC, quotas, audit logging
  • Cloud integration - Replicate, self-hosted workers
  • Memory optimization - DeepSpeed, FSDP2, quantization

Supported Models

Type Models
Image Flux.½, SD3, SDXL, Chroma, Auraflow, PixArt, Sana, Lumina2, HiDream, and more
Video Wan, LTX Video, Hunyuan Video, Kandinsky 5, LongCat
Audio ACE-Step

See Model Guides for complete documentation.

Community

License

SimpleTuner is open source software.