Roadmap discussion: what should v0.2 look like? #39
Replies: 3 comments
|
Here is where I think v0.2 should focus, roughly in priority order: 1. SDXL backbone upgrade. This is probably the single biggest quality improvement we can make. SD1.5 caps us at 512x512 native resolution, which is limiting for clinical use. SDXL at 1024x1024 would make the output much more convincing. The ControlNet conditioning pipeline should port over without major changes since there are good SDXL ControlNet checkpoints available now. 2. IP-Adapter FaceID. Identity drift is the number one complaint in practice. The current pipeline sometimes shifts skin tone, eye color, or subtle facial proportions even with the histogram matching and Laplacian blending in post-processing. IP-Adapter FaceID gives us an identity embedding that gets injected into the cross-attention layers, which should lock down identity much better than post-hoc corrections. 3. LoRA fine-tuning. This is where things get really practical. A surgeon who does a specific style of rhinoplasty (say, preservation rhinoplasty vs. open approach) could fine-tune on their own before/after pairs. Even 20-50 pairs with a LoRA rank of 8-16 should give meaningfully different results. 4. Data-driven displacement model. We already have the infrastructure for this (the DisplacementModel class), and the Rathgeb et al. database gives us real surgical pairs to fit against. This would replace the hand-tuned RBF parameters with learned displacement fields per procedure, which should produce more realistic deformations. The HF Spaces demo is already live (CPU/TPS mode), so that one is checked off. More procedure presets are ongoing; see Discussion #38. What would make you use this in practice? Curious especially whether the SDXL upgrade or the identity preservation matters more to people testing this out. |
|
"Identity preservation is the bigger blocker for me in practice. SDXL is a quality upgrade but IP-Adapter FaceID solves an actual correctness problem — if the patient's skin tone or eye color shifts, the simulation loses clinical credibility entirely. I'd prioritize FaceID over SDXL. |
|
Solid points, @P-r-e-m-i-u-m. I agree that identity preservation is the more fundamental problem. SDXL gives us higher resolution, but if the patient does not look like themselves in the output, the resolution is irrelevant clinically. The sequencing you suggest makes sense: solve identity drift with IP-Adapter FaceID first, then upgrade the backbone to SDXL, then layer LoRA on top of a pipeline that already preserves identity reliably. On data-driven displacements: glad to hear the hand-tuned RBF iteration was nontrivial. That validates the priority of fitting displacement fields from real surgical pairs (the Rathgeb et al. database). Once we have learned fields, community presets like your mentoplasty contribution become even more powerful because they can be validated against real surgical outcomes rather than just anatomical heuristics. Updating the v0.2 priorities to reflect this: FaceID first, SDXL second. |
Uh oh!
There was an error while loading. Please reload this page.
v0.1 shipped with the core pipeline (MediaPipe landmarks, Gaussian RBF deformation, ControlNet conditioning, TPS/img2img/controlnet inference).
For v0.2, we are considering:
What features matter most to you? What would make you use this in practice?
Drop your thoughts below.
All reactions