Real run · self-run demonstration

Which model version is better — recraftv4 vs recraftv4_1.

A real evaluation we ran ourselves: five prompts, both model versions, scored on four dimensions. This is the deliverable format exactly as a client would receive it — on real generated images, not mock numbers.

Honest method. 10 images (5 prompts × 2 versions) generated at 1024×1024 with Recraft, then scored 0–100 per dimension by an automated visual judge. A production MetronLabs evaluation adds a blind human panel, inter-rater agreement and confidence intervals — the numbers below are single-judge and directional.

Result — score by dimension

Dimensionrecraftv4 (A)recraftv4_1 (B)Δ
Prompt adherence86.887.6+0.8
Anatomy85.587.5+2.0
Text rendering75.079.5+4.5
Aesthetics83.687.8+4.2
Verdict: recraftv4_1 improves on every dimension. Biggest gains in aesthetics (+4.2) and text rendering (+4.5); prompt adherence is essentially flat (+0.8, within noise). If you're upgrading, v4.1 is the safe move — the one thing to keep watching is prompt adherence, where the gap is not yet meaningful.

The images that were scored

recraftv4 vs recraftv4_1 across five prompts

Prompts & what each tests

Want this run on your own model or agent?