Sample design
The planned baseline contains 100 licensed or internally authorized images. Each image is assigned to one primary category before testing, deduplicated by file hash and run once in randomized order. A 20-image subset is rerun to inspect repeatability; those repeat runs are reported separately and never added to the headline success-rate denominator.
| Category | Images | Coverage |
|---|---|---|
| People and crowds | 20 | Hair, limbs, overlapping subjects, partial occlusion |
| Product photos | 20 | Plain and contextual backgrounds, reflections, packaging |
| Complex textures | 20 | Grass, brick, water, fabric, repeated patterns |
| Small distractions | 20 | Wires, signs, litter, logos and edge objects |
| Hard cases | 20 | Large occlusions, shadows, text and ambiguous instructions |
Conditions recorded for every run
Every raw record must include the source reference, display authorization, SHA-256 file hash, width, height, byte size, format, exact prompt, UTC timestamp, test region, browser or API client, model identifier returned by the service, task ID, credit cost, status code and result dimensions.
- No silent replacement of failed images or prompts.
- The same production endpoint, credit rules and storage path used by customers.
- Warm-up requests are labeled and excluded before the test starts.
- Timeout, provider, storage and credit errors remain in the raw log.
- Raw images are not made public unless the authorization explicitly permits it.
Metrics and current status
Sample count
100 planned test imagesA fixed, deduplicated set: 40 images with longest side ≤1024px, 40 at 1025–2048px, and 20 above 2048px. Format target: 40 JPG, 30 PNG, 30 WebP.
Dimension retention rate
Unknown — baseline not runSuccessful outputs whose pixel width and height both equal the source, divided by all successful outputs. A changed file format does not count as a dimension change.
p50 / p95 duration
Unknown — baseline not runWall-clock duration from accepted POST request to the successful API response. Reported in seconds by test region, with p50 as the median and p95 as the 95th percentile.
Success rate
Unknown — baseline not runTasks that return a readable, downloadable result and are recorded as SUCCESS, divided by all submitted tasks. Retries keep the same task ID and do not inflate the denominator.
Reporting rule: every number must be labeled Measured, include the observation date and sample count, and link to a versioned test summary. Estimates and proxy measurements cannot be presented as test results.
Known limitations
Object removal is generative editing, so an API success does not guarantee a visually acceptable edit. Large occlusions, faces and hands, text, logos, transparent objects, reflections, cast shadows, repeating patterns and objects touching the frame can produce reconstruction errors. Output dimensions or format may differ from the source. Results should be reviewed before publication or commercial use.
The baseline represents the declared sample and environment only. It does not predict performance for every image, location, network condition or future model version. Model, prompt, provider or pipeline changes require a new test version rather than silently replacing the old result.
Quarterly review and change control
On the first business week of January, April, July and October, the maintainer reviews pricing, active model identifiers, third-party processors and retention behavior against production configuration and provider records. Material changes update the relevant public page and create a dated entry in the review log.
See the Privacy Policy, credit rules and case-study publication standard for the current customer-facing controls.