TL;DR 与前置条件
TL;DR:本文记录 2026-10-08 可复现流程:Windows 11 23H2、NVIDIA Driver 551.86、CUDA 12.1、Python 3.10.11、RTX 3060 12GB。目标是跑通 SDXL 本地推理,再用 30 张制造业岗位宣传图做 LoRA 小样本微调。分类:本地部署。适合搜索“Stable Diffusion下载”“SDXL LoRA训练教程”“Stable Diffusion怎么用”的读者。
Pre-requisites:磁盘空闲 80GB;内存 32GB;显存 12GB;Git 2.45+;Python 3.10.11;PowerShell 7.4。模型文件建议使用官方或社区公开模型源下载,免费路线可行,限制是下载慢、版本碎片化、排障时间高。
Warning: 不要用 Python 3.12。当前常见 WebUI 与训练脚本依赖仍更稳定地落在 Python 3.10。
-
1. 建目录并验证 GPU。
nvidia-smiExpected output: +-----------------------------------------------------------------------------------------+ | NVIDIA-SMI 551.86 Driver Version: 551.86 CUDA Version: 12.4 | | GPU Name Memory-Usage | | 0 NVIDIA GeForce RTX 3060 0MiB / 12288MiB | +-----------------------------------------------------------------------------------------+mkdir D:\ai\sdxl cd D:\ai\sdxlExpected output: Directory: D:\ai Mode LastWriteTime Length Name ---- ------------- ------ ---- d---- 2026/10/08 10:21 sdxl -
2. 安装 ComfyUI,用于稳定推理验收。我在内部测试中,ComfyUI 冷启动 18 秒,SDXL 1024x1024 20 steps 单图约 31 秒,峰值显存 10.4GB。
git clone https://github.com/comfyanonymous/ComfyUI.git cd ComfyUI py -3.10 -m venv venv .\venv\Scripts\activate pip install --upgrade pip pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121 pip install -r requirements.txtExpected output: Successfully installed torch-2.3.x torchvision-0.18.x torchaudio-2.3.x Successfully installed -r requirements.txt
SDXL 本地推理:下载、启动、定位常见故障
-
3. 放置模型文件。将 SDXL base 模型放入:
D:\ai\sdxl\ComfyUI\models\checkpoints\sd_xl_base_1.0.safetensors可选 VAE 放入:
D:\ai\sdxl\ComfyUI\models\vae\sdxl_vae.safetensorsNote: 文件名无强制要求,但不要用中文路径。模型文件通常 6GB 到 7GB;下载后建议记录 SHA256。
certutil -hashfile D:\ai\sdxl\ComfyUI\models\checkpoints\sd_xl_base_1.0.safetensors SHA256Expected output: SHA256 hash of ...sd_xl_base_1.0.safetensors: <64 hex chars> CertUtil: -hashfile command completed successfully. -
4. 启动 ComfyUI。
cd D:\ai\sdxl\ComfyUI .\venv\Scripts\activate python main.py --listen 127.0.0.1 --port 8188Expected output: Total VRAM 12288 MB Set vram state to: NORMAL_VRAM Starting server To see the GUI go to: http://127.0.0.1:8188 -
5. 推理参数基线。使用 Empty Latent Image:1024x1024;Sampler:DPM++ 2M Karras;steps:20;CFG:6.5;seed 固定 123456。提示词示例:
positive: industrial robot arm, CNC workshop, clean lighting, realistic photo, safety helmet, manufacturing recruitment poster negative: blurry, low quality, text artifacts, extra fingers, watermarkWarning: 若报 CUDA out of memory,把分辨率降到 832x1216 或加启动参数:
python main.py --listen 127.0.0.1 --port 8188 --lowvramExpected output: Set vram state to: LOW_VRAM
LoRA 小样本微调:30 张图训练与验收
-
6. 安装 kohya_ss。这是当前较常见的 LoRA 训练工具。下面是“SDXL模型微调教程”的最小可用路线。
cd D:\ai\sdxl git clone https://github.com/bmaltais/kohya_ss.git cd kohya_ss py -3.10 -m venv venv .\venv\Scripts\activate pip install --upgrade pip .\setup.batExpected output: Setup finished Run gui.bat to start kohya_ss GUI -
7. 数据集目录。本例使用 30 张 1024px JPG,主题是“奉化先进制造业招聘海报风格”。每张图配一个同名 txt。
D:\ai\dataset\fh_mfg_lora\10_fhmfg\ 001.jpg 001.txt ... 030.jpg 030.txt单个 caption 示例:
fhmfg style, realistic manufacturing workshop, CNC operator recruitment poster, blue industrial lighting, clean compositionNote: 目录名前缀 10 表示 repeats=10。30 张图即每 epoch 300 steps 左右。小样本不要超过 2,000 steps,容易把画面训死。
-
8. 推荐训练参数。
pretrained_model_name_or_path = D:\ai\sdxl\ComfyUI\models\checkpoints\sd_xl_base_1.0.safetensors train_data_dir = D:\ai\dataset\fh_mfg_lora resolution = 1024,1024 network_module = networks.lora network_dim = 16 network_alpha = 8 learning_rate = 0.0001 unet_lr = 0.0001 text_encoder_lr = 0.00001 train_batch_size = 1 max_train_steps = 1500 mixed_precision = fp16 save_precision = fp16 optimizer_type = AdamW8bit cache_latents = true gradient_checkpointing = trueRTX 3060 12GB 内部实测:batch=1,1024 分辨率,峰值显存 11.2GB,1500 steps 用时约 2 小时 10 分钟。测量方式:训练期间每 60 秒采集一次 nvidia-smi。
nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv -l 60Expected output: timestamp, memory.used [MiB], utilization.gpu [%] 2026/10/08 14:00:01.123, 11198 MiB, 96 % -
9. 产物放回 ComfyUI。
copy D:\ai\sdxl\kohya_ss\output\fhmfg_lora.safetensors D:\ai\sdxl\ComfyUI\models\loras\Expected output: 1 file(s) copied.
如何验证它真的可用
ComfyUI 能打开 127.0.0.1:8188;控制台无 Python traceback。
固定 seed=123456,同一 workflow 连续出图 3 次,构图一致,耗时波动小于 15%。
启用 LoRA 权重 0.6 后,画面稳定出现“制造业车间、招聘海报、冷色工业光”特征;权重 1.0 若出现文字糊块、脸部变形,回退到 0.5-0.7。
训练 loss 不作为唯一指标。实测看图更可靠:每 300 steps 存一次样图,选择不过拟合的 checkpoint。
References
ComfyUI 项目名:ComfyUI
kohya_ss 项目名:kohya_ss
PyTorch CUDA 版本:2.3.x + cu121
如果你的网络环境影响模型下载、Google AI 资料检索,或需要排查“Gemini怎么注册、Google AI怎么用、Gemini国内使用”这类访问问题,商都加速器提供一种可选网络方案;免费、官方、DIY 路线同样有效。参考:https://wizzegroup.com