Human-Aligned Evaluation of Text-to-SVG Generation
1 Hasso Plattner Institute, Germany · 2 University of Modena and Reggio Emilia, Italy
† Equal contribution.
Abstract
SVG generation is attracting increasing attention as generative models improve in expressiveness and controllability. Progress, however, is held back by the lack of domain-specific evaluation protocols: current practice relies on metrics designed for natural images, most notably CLIPScore, which was never trained on vector graphics and aligns only partially with human judgment.
We introduce SVG-Score, a human-aligned evaluation framework for text-to-SVG generation. Through controlled caption and image perturbations, we first show that CLIP-based scores barely react to the errors SVG generators actually make — wrong colors, counts and spatial relations — and that off-the-shelf VLM judges, while more sensitive, respond unevenly across error types and SVG styles. We then introduce a human-annotated dataset for Semantic Alignment, measuring how faithfully a generated SVG reflects its caption, and build two complementary evaluators on it: CLIP scorers adapted to vector graphics and aligned to human preferences, for fast large-scale evaluation, and a VLM judge trained with supervised fine-tuning and reward-shaped reinforcement learning, for more expressive and interpretable assessment. Using both, we benchmark major open-source, commercial and optimization-based SVG generators on an independent caption set.
The problem
We take one SVG and the caption that belongs to it. We change the colour, the position or the count in the caption, and CLIPScore barely moves. We throw the drawing away instead, and it falls by 14 points.
OmniSVG sample · original SVG
Moving the splashes below the whale costs 0.32 points. Turning four splashes into seven costs 0.19. Only colour registers, at 3.32. A plain blue square costs 13.75. CLIP sees that the picture changed, not that the caption became false.
Our answer
We label caption–SVG pairs by hand and train four evaluators on them. Below is how well each one agrees with our annotators on the held-out test split (Spearman ρ ×100).
8,671 SVGs and 1,858 captions, each pair rated by one of five annotators. A 1 means the SVG is unrelated to the caption, a 5 means it is a faithful rendering of it.
We fine-tune CLIP on 3.2M SVG–caption pairs, then train LoRA adapters on the human preferences. Scoring is still one forward pass, so it runs over a whole benchmark cheaply.
Qwen3-VL-8B, taught to write a short rationale and a score, then trained with GRPO on two rewards. It agrees with our annotators more closely than any judge we tried, open or commercial.
Analysis
We built five perturbations and ran them on 200 samples. Three edit the caption and keep the SVG; two keep the caption and replace the SVG. An evaluator that reads meaning should score lower in all five cases.
Each mentioned colour is replaced by an alternative; averaged over five variants.
Spatial and relational terms are replaced by their semantic opposite.
Each quantity q becomes q−1, q+1 and one larger perturbation.
The SVG is replaced by one coloured circle per colour named in the caption.
The SVG is replaced by a uniform canvas in one colour named in the caption.
Left: the mean score change when a pair is corrupted. CLIP is in CLIPScore points, the VLM on the 1–5 scale, so the two are not comparable. Right: how much further each of our evaluators drops than its own baseline. Our CLIP mainly gets better at noticing the picture is gone; our VLM judge gains on all three caption edits.
Method
We trained three CLIP scorers and one VLM judge on the same human ratings. Use CLIP as a fast in-training signal, cheap enough to score every checkpoint. Use the VLM for the exhaustive final evaluation, where it also explains each score.
Off-the-shelf checkpoints, trained on natural images.
Drawn from StarVector and OmniSVG with the overlap removed. The original captions are noisy, so we recaption every image with Qwen3-VL-8B, then fine-tune both encoders contrastively.
For each caption we pair two SVGs with different human scores and treat the higher-rated one as preferred. The base weights stay frozen; only the adapters learn.
Already a better starting point than any CLIP scorer, but its scores are not calibrated.
One epoch of supervised fine-tuning to produce
<think>…</think><score>1–5</score>. On its own this improves
calibration but lowers rank correlation.
Two completions per input, with SVGs that share a caption kept in the same batch.
Dataset 1
We pair every caption with seven SVGs of differing quality, found by retrieval from an OmniSVG sample. Five annotators rate each pair from 1 to 5. That gives, for every caption, an ordering the evaluators learn to reproduce.
| Split | Ratings | SVGs | Captions |
|---|---|---|---|
| Train | 10,583 | 6,312 | 1,519 |
| Test | 2,374 | 2,359 | 339 |
Dataset 2
We wrote 1,616 captions at three levels of difficulty and checked them by hand. We generated one SVG per caption with 16 models. None of our evaluators was trained on this set.
These are the generators' own SVG files, drawn by your browser rather than screenshotted. A cell marked no output returned nothing we could parse.
Results
All numbers are on the held-out test split. PA is pairwise accuracy: how often an evaluator orders two SVGs of the same caption the way the humans did. Correlations ×100, higher is better except MAE.
| Judge | Params | ρ ↑ | r ↑ | τ ↑ | MAE ↓ | PA ↑ |
|---|---|---|---|---|---|---|
| Natural-image preference scorers | ||||||
| Aesthetic | 304.9M | 17.33 | 23.64 | 12.68 | 1.66 | 57.66 |
| ImageReward | 446.6M | 56.95 | 55.83 | 43.27 | 1.08 | 75.66 |
| PickScore | 986.1M | 45.63 | 36.92 | 34.61 | 1.42 | 75.72 |
| HPSv2 | 986.1M | 55.23 | 53.91 | 42.19 | 1.26 | 75.66 |
| CLIP ViT-B/32 | ||||||
| Vanilla | 151.3M | 42.90 | 41.10 | 32.36 | 1.43 | 69.35 |
| + SVG fine-tuned | 151.3M | 54.26 | 54.44 | 41.41 | 1.27 | 75.82 |
| SVG-Score (ours) | 151.3M | 59.31 | 58.43 | 45.46 | 1.18 | 77.87 |
| CLIP ViT-L/14 | ||||||
| Vanilla | 427.6M | 49.80 | 49.23 | 37.81 | 1.33 | 72.75 |
| + SVG fine-tuned | 427.6M | 54.40 | 54.43 | 41.42 | 1.27 | 76.39 |
| SVG-Score (ours) | 427.6M | 60.87 | 59.69 | 46.88 | 1.18 | 78.37 |
| CLIP ViT-H/14 | ||||||
| Vanilla | 986.1M | 56.64 | 56.79 | 43.73 | 1.21 | 76.87 |
| + SVG fine-tuned | 986.1M | 59.72 | 58.59 | 45.71 | 1.22 | 77.52 |
| SVG-Score (ours) | 986.1M | 63.18 | 58.95 | 48.62 | 1.15 | 80.64 |
Adapting to the SVG domain helps on every backbone, and preference alignment adds more on top. The gain is largest on the smallest backbone, +16.41 ρ on ViT-B/32 against +6.54 on ViT-H/14: a stronger visual prior already absorbs part of the domain shift. Our ViT-H/14 beats HPSv2 by 7.95 ρ at the same parameter count.
| Evaluator | ρ ↑ | r ↑ | τ ↑ | MAE ↓ | PA ↑ |
|---|---|---|---|---|---|
| Open models · zero-shot | |||||
| LLaVA-1.5-7B | 18.30 | 22.69 | 16.08 | 1.48 | 56.00 |
| Qwen3.5-9B | 50.34 | 49.74 | 43.28 | 1.04 | 69.14 |
| InternVL3-8B | 60.38 | 61.03 | 52.14 | 0.95 | 69.54 |
| Gemma-3-12B-it | 66.70 | 65.82 | 57.86 | 0.89 | 72.47 |
| Qwen3-VL-8B | 67.76 | 67.48 | 59.14 | 0.80 | 74.97 |
| Closed models · zero-shot | |||||
| GPT-5.4-nano | 47.61 | 48.29 | 40.68 | 1.09 | 69.58 |
| GPT-5.6-luna | 63.76 | 64.02 | 55.62 | 0.91 | 76.46 |
| GPT-5.4-mini | 66.96 | 67.22 | 58.58 | 0.84 | 78.43 |
| Claude Haiku 4.5 | 67.12 | 66.86 | 57.92 | 0.82 | 77.87 |
| SVG-specific | |||||
| VectorGym | 65.12 | 61.93 | 56.88 | 0.99 | 75.38 |
| SVG-Score (ours) | 74.85 | 74.90 | 65.08 | 0.68 | 79.59 |
VectorGym is the only earlier evaluator trained on SVG tasks, and it still lands below four off-the-shelf VLMs: being good at SVG tasks does not by itself mean agreeing with human ratings. Our judge improves on its own zero-shot backbone in every column. The margin is narrowest on pairwise accuracy, so the training buys absolute calibration more than relative ordering.
| SFT | GRPO | Ordinal | Ranking | ρ ↑ | r ↑ | τ ↑ | MAE ↓ | PA ↑ |
|---|---|---|---|---|---|---|---|---|
| — | — | — | — | 67.76 | 67.48 | 59.14 | 0.80 | 74.97 |
| ✓ | — | — | — | 66.40 | 66.13 | 56.44 | 0.73 | 76.67 |
| — | ✓ | ✓ | — | 69.32 | 69.16 | 60.50 | 0.83 | 74.44 |
| ✓ | ✓ | ✓ | — | 72.95 | 72.98 | 62.88 | 0.69 | 78.68 |
| ✓ | ✓ | ✓ | ✓ | 74.85 | 74.90 | 65.08 | 0.68 | 79.59 |
Neither stage is enough on its own. SFT alone lowers rank correlation while improving calibration; GRPO alone raises ρ but pushes MAE past the zero-shot baseline. Together they reach 72.95 ρ: the policy needs supervised imitation to start from a well-formed output distribution. Adding the intra-caption ranking reward then lifts ρ and τ and leaves MAE where it was, which is what a reward defined on relative scores should do.
| Generator | SVG-CLIP-B ↑ | SVG-CLIP-L ↑ | SVG-CLIP-H ↑ | HPSv2 ↑ | Aesthetic ↑ | Qwen ZS ↑ | Ours VLM ↑ | Error % ↓ |
|---|---|---|---|---|---|---|---|---|
| Painterly rendering | ||||||||
| CLIPDraw | 31.81 | 18.30 | 18.71 | 17.06 | 4.10 | 1.68 | 1.33 | 0.00 |
| VectorFusion | 27.88 | 20.89 | 22.03 | 19.27 | 4.59 | 2.50 | 2.09 | 0.00 |
| SVGDreamer | 24.31 | 17.39 | 18.96 | 18.46 | 4.61 | 1.83 | 1.50 | 0.00 |
| DiffSketcher | 19.75 | 14.11 | 15.86 | 13.08 | 4.05 | 1.01 | 1.00 | 4.39 |
| Commercial | ||||||||
| GPT-5 nano | 31.57 | 27.04 | 35.16 | 20.06 | 4.40 | 3.24 | 2.85 | 1.67 |
| GPT-4o | 33.25 | 27.20 | 27.16 | 19.76 | 4.39 | 2.85 | 2.49 | 0.74 |
| GPT-5 mini | 36.66 | 31.13 | 30.61 | 21.63 | 4.58 | 3.90 | 3.61 | 1.30 |
| Claude Sonnet 5 | 38.44 | 32.91 | 31.94 | 23.08 | 4.67 | 4.13 | 3.99 | 0.25 |
| Gemini 3.0 Flash | 32.16 | 28.24 | 30.97 | 22.47 | 4.56 | 3.56 | 3.42 | 2.29 |
| Open | ||||||||
| LLM4SVG | 20.56 | 16.03 | 18.53 | 11.36 | 3.24 | 1.40 | 1.09 | 24.57 |
| vHector | 24.30 | 19.53 | 21.98 | 14.42 | 3.92 | 1.47 | 1.24 | 12.07 |
| OmniSVG | 27.79 | 22.53 | 26.63 | 16.90 | 4.49 | 2.08 | 1.74 | 0.31 |
| SVGen | 20.63 | 16.51 | 19.75 | 13.25 | 3.54 | 1.77 | 1.29 | 23.45 |
| IntroSVG | 18.94 | 15.62 | 19.39 | 12.36 | 3.07 | 2.25 | 1.75 | 33.17 |
| InternSVG | 28.94 | 22.75 | 23.10 | 18.10 | 4.48 | 2.08 | 1.73 | 0.31 |
| HiVG | 29.98 | 25.14 | 31.06 | 18.57 | 4.43 | 2.21 | 1.86 | 0.37 |
Commercial systems lead, with Claude Sonnet 5 first under six of the seven evaluators. Among the rest, HiVG leads the open group under all three SVG-CLIP variants and under our judge, while the painterly VectorFusion takes the highest VLM score of any non-commercial system. Reliability varies a lot: up to 33.17% of IntroSVG's outputs fail to parse. Failures get the minimum evaluator score.
Citation
We will update this once the preprint is online.
@article{cipriano2026svgscore,
title = {SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation},
author = {Cipriano, Marco and Zini, Leonardo and Schild, Alexandra and
Teutschbein, Valentin and Mimi, Afsana and Cornia, Marcella and
Baraldi, Lorenzo and de Melo, Gerard},
journal = {arXiv preprint},
year = {2026}
}