arxiv:2609.20633
Published on Sep 17
· Submitted by
Yulong Chen on Sep 21
· City University of Hong Kong
Upvote
2
Authors:
,
,
,
,
Abstract
Text-guided image editing must introduce the requested changes while preserving unrelated source content. Diffusion-based editors rely on spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal autoregressive editors face a further constraint: their fixed decoding order limits revision of earlier decisions. We introduce RefineEdit, a training-free prompt-to-prompt image editing framework built on a Generative Refinement Network. Our key idea is to couple edit localization with content generation through the global refinement of binary image codes, allowing editing evidence to be reassessed as the image evolves. RefineEdit initializes an editing branch from an intermediate source state, reusing the emerging layout. We compare the probabilities assigned by the two branches to the same source-sampled bits, using their signed differences to select editable positions and bits. Selected bits follow editing refinement, while the remaining bits copy the evolving source state. To stabilize editing across refinement steps, adaptive spatial freezing limits unnecessary mask expansion, while finite bit locking keeps recently selected bits editable. The framework requires no additional training, external masks, or attention control. Across nine editing categories of PIE-Bench, RefineEdit achieves the best background-preservation scores in PSNR, LPIPS, MSE and SSIM, together with the highest whole-image and edited-region CLIP scores among the evaluated methods.
View arXiv page View PDF GitHub 7 Add to collection
Community
Paper author Paper submitter 1 day ago
about 13 hours ago
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
-
Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective (2026)
-
Diffusion Image Editing via Asynchronous Token Decoding (2026)
-
RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing (2026)
-
SR-Edit: Region-Aware Image Editing via Self-Refinement (2026)
-
MaskFlow: Precise, Consistent and Seamless Regional Image Editing (2026)
-
One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing (2026)
-
EditFlow3D: Automated Local Editing of 3D Assets with Trajectory Preservation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on HF Mirror checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
· Sign up or log in to comment
Upvote
2
Get this paper in your agent:
hf papers read 2609.20633
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2609.20633 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2609.20633 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2609.20633 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.