Mina LabsMINA LABS Start creating free
Blog / News
GPT Image 2.5 Flare Gets a Forgery Task Test

GPT Image 2.5 Flare Gets a Forgery Task Test

2026-09-15

A new arXiv paper tests whether the improvements associated with ChatGPT Images 2.5 translate into better results on image forgery tasks with known answers. The study compares the Flare and Sunburst API models with GPT-Image-2, using the same week of testing. Read the full paper, “From Advertised Improvements to Measured Capabilities”.

The evaluation covers four tasks: changing fields on receipts, making repeated edits, placing products into images, and rendering fine print. The researchers compare the generated images after registration, then use OCR and task-specific checks to measure what changed and what remained intact.

The useful part of this study is its focus on measurable behavior rather than general image quality. A model can produce an image that looks convincing at a glance while still changing nearby text, dropping earlier edits, or rendering product codes incorrectly. Those details matter when the image is part of a mockup, catalog, advertisement, packaging concept, or other production workflow.

On receipt-field edits, Flare and Sunburst changed less surrounding receipt text than the GPT-Image-2 baselines after the images were aligned. The paper reports OCR-detected changes to nearby text at 31.7% for Flare and 31.2% for Sunburst, compared with 44.2% for both GPT-Image-2 baselines. The improvement appeared mainly on CORD receipts.

That result comes with an important limit. The lower rate of surrounding changes did not come with a detectable improvement in getting the requested target field correct. In practical terms, a model may preserve more of the surrounding document while still requiring review of the specific value or field that was meant to change.

Repeated editing produced a similar split. Flare retained fewer earlier edits on CORD receipts, according to the paper. Photo-editing sequences showed little separation between the models. This suggests that the benefit of using one model over another can depend on the source image and the type of edit sequence, rather than appearing consistently across every workflow.

The product-placement results were different. Product codes were more often legible with Images 2.5, while the inserted products were also placed at a larger scale. The researchers note that their analysis does not establish a fidelity gain independent of that larger product placement. That distinction matters when judging whether the model is actually preserving detail or simply making the object easier to read by displaying it more prominently.

What it costs and where to use it

GPT Image 2.5 Flare is available on Mina Labs. A single text-to-image generation costs 2, and a single edit costs 2. The model can therefore fit into workflows that need quick visual iterations, provided each generation is reviewed for text accuracy, retained edits, and layout changes.

For makers, the main practical value is not that the paper proves Flare is better in every category. It does not. Instead, the results identify areas worth testing in your own pipeline. If you work with product mockups, receipt-like layouts, packaging studies, or images that combine objects with small text, Flare gives you a model to evaluate against those specific requirements.

What we would use it for

We would use Flare for product-placement concepts where the inserted object needs to remain readable and visually prominent. This includes early advertising layouts, ecommerce mockups, retail scenes, and packaging explorations. We would generate several variations, then check product codes and small labels manually before using any output beyond the concept stage.

We would also use it for controlled image edits where preserving the surrounding composition is important. The receipt results suggest that Flare may reduce some unwanted changes to nearby text on certain document images. That makes it worth testing for localized edits, but not treating the output as a guaranteed pixel-preserving transformation.

For repeated edits, we would keep the workflow short and save each approved version. The study’s finding about retained earlier edits is a reminder that an edit sequence can drift. Instead of stacking many changes onto one image, we would compare checkpoints and return to the cleanest approved version when a later instruction damages an earlier one.

Fine print would remain a review-heavy use case. Better legibility is useful for concepts, but readable text is not necessarily accurate text. We would use the model to explore hierarchy, placement, and visual direction, then replace or verify important copy in a separate production step.

The paper offers a practical way to think about GPT Image 2.5 Flare: test it by task, not by reputation. Its results show fewer surrounding receipt changes in one part of the evaluation, weaker separation in photo-edit sequences, and more legible product codes alongside larger placements. Those are useful signals for deciding where to try it today, while leaving room for human checking wherever exact text and edit fidelity matter.

Source: arXiv cs.CV, https://arxiv.org/abs/2609.13617 Make something with itMina Labs runs these models in your browser. Pay per generation, no subscription.