From ca484a7a90c81a7682c92259d0c41e6aea80c842 Mon Sep 17 00:00:00 2001 From: cuhawk Date: Mon, 31 Aug 2026 18:39:45 +0530 Subject: [PATCH] fix: use the offline mask blur ratio in the realtime blending path get_image() (offline) and get_image_prepare_material() (realtime) build the blending mask the same way: face_seg over the expanded crop, crop to the face box, paste onto a black canvas, drop everything above upper_boundary_ratio, then Gaussian blur to soften the edge. The blur is the only step where the two disagree. get_image() sizes the kernel at 0.05 of the crop width and get_image_prepare_material() at 0.1, so the realtime mask feathers over twice the distance. With the default expand=1.5 a 149x200 face gives a 300 px crop, and the two kernels come out at 15 px and 31 px. Where the mask edge runs close to the mouth, the wider ramp mixes more of the original frame back over the generated pixels than the offline path would from the same inputs. What the file history shows, because it cuts both ways: - Before v1.5 both functions used 0.1 and agreed. - v1.5 (db20431) rewrote get_image only: expand 1.2 -> 1.5, new mode and fp arguments, blur ratio 0.1 -> 0.05. It left get_image_prepare_material untouched on the older code. - 39ccf69 ("feat: real-time infer", #286) then ported that rewrite into get_image_prepare_material: expand 1.2 -> 1.5, added fp and mode, passed both through to face_seg. It edited the function line by line and left the 0.1 in place. So this is not a value that only one commit overlooked; a later commit went through the same function and kept it. Nothing in the code or the commit messages argues for it either way. The reason it still reads as a leftover is that 39ccf69 carried over every other difference the v1.5 rewrite had introduced here, and the blur ratio is the only one it did not. If the wider feather is deliberate for the realtime path, please close this: a comment recording why would be worth more than the change. Testing note: scripts/realtime_inference.py writes every computed mask to {avatar_path}/mask/*.png and pickles the crop boxes, and later runs read those back instead of recomputing them. get_image_prepare_material() is called only from prepare_material(), so this change is inert for an avatar that has already been prepared. Reproducing any difference needs a fresh avatar_id, or the existing avatar directory removed first. --- musetalk/utils/blending.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/musetalk/utils/blending.py b/musetalk/utils/blending.py index fa3effcd..476aa2ce 100755 --- a/musetalk/utils/blending.py +++ b/musetalk/utils/blending.py @@ -131,6 +131,6 @@ def get_image_prepare_material(image, face_box, upper_boundary_ratio=0.5, expand modified_mask_image = Image.new('L', ori_shape, 0) modified_mask_image.paste(mask_image.crop((0, top_boundary, width, height)), (0, top_boundary)) - blur_kernel_size = int(0.1 * ori_shape[0] // 2 * 2) + 1 + blur_kernel_size = int(0.05 * ori_shape[0] // 2 * 2) + 1 mask_array = cv2.GaussianBlur(np.array(modified_mask_image), (blur_kernel_size, blur_kernel_size), 0) return mask_array, crop_box