VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
The First General Reward Model for Joint Video-Audio Generation.
Abstract
Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions separately and fail to capture the overall semantic and temporal coherence among the text prompt, video, and audio that shapes human preferences. More critically, they are poorly aligned with actual human judgments. Optimizing models against these metrics encourages reward hacking, generating video-audio content that achieves high scores on these metrics yet appears incoherent or unfaithful to human viewers. To address this problem, we first construct VAPref-10K for joint video-audio generation, comprising 9K prompts and 10.3K fine-grained paired comparisons from generation models. We also introduce the VA-Judger-Bench benchmark with both in-domain and out-of-domain model comparisons to evaluate whether reward models truly align with human preferences. We further propose VA-Judger, a chain-of-thought omni-reward model for joint video-audio generation. VA-Judger learns from clear quality gaps, distills format-valid responses for harder near-quality comparisons through rejection sampling verified against human annotations, and performs dimension-wise reinforcement learning for denser reward signals. Experiments show that VA-Judger outperforms metric baselines in predicting human preferences on both in-domain and out-of-domain evaluations. Using its human-aligned rewards to post-train an audio-video generation model also improves most automatic metrics and human preference.
Pipeline

VA-Judger training pipeline. The model learns the comparison rubric from easy pairs, aligns with human feedback on hard pairs, and is refined with Dimension-Wise GRPO.
Performance
VA-Judger Performance on VA-Judger-Bench
| Model | Easy | In-domain | Out-of-domain | Overall |
|---|---|---|---|---|
| single-dimension evaluation metrics | ||||
| Video quality: VideoAlign | 62.25 | 53.20 | 48.40 | 54.26 |
| Audio quality: Audiobox Aesthetics | 58.50 | 50.00 | 44.00 | 50.35 |
| Text-video: CLIP Score | 55.75 | 54.80 | 53.80 | 54.70 |
| Text-audio: ImageBind T-A | 58.50 | 52.80 | 51.60 | 54.26 |
| Audio-video: ImageBind | 61.25 | 55.60 | 51.80 | 55.91 |
| Synchronization: SynchFormer | 61.50 | 46.80 | 48.20 | 52.52 |
| Overall: Javis Score | 60.00 | 55.60 | 55.40 | 57.04 |
| Metric ensemble: Hard voting | 64.50 | 52.00 | 50.00 | 55.48 |
| Metric ensemble: Z-score soft voting | 67.00 | 58.00 | 52.40 | 58.70 |
| Reward models | ||||
| Qwen3-Omni Captioner (No CoT) | 57.00 / 57.00 | 54.00 / 54.00 | 56.20 / 56.20 | 56.00 / 56.00 |
| Qwen3-Omni Captioner (CoT) | 52.50 / 57.69 | 49.60 / 54.87 | 46.60 / 56.97 | 49.30 / 56.76 |
| Qwen3-Omni Instruct NoCoT | 60.75 / 60.75 | 54.00 / 54.00 | 54.60 / 54.60 | 56.61 / 56.61 |
| Qwen3-Omni Instruct CoT | 63.25 / 64.71 | 54.80 / 55.92 | 55.00 / 55.33 | 57.83 / 58.69 |
| VA-Judger (Ours) | 76.25 / 76.25 | 66.00 / 66.00 | 63.40 / 63.40 | 68.43 / 68.43 |
Table 1. Accuracy (%) on 400 easy, 250 in-domain, and 500 out-of-domain pairs. Reward models report Total Acc / Parsed Acc, and Overall is micro accuracy over all 1,150 pairs.
Video Generation Model Performance on JavisBench
| Video Quality | Audiobox Aesthetics | Cross Modal Alignment | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | VQ ↑ | MQ ↑ | AQ ↑ | AB CE ↑ | AB CU ↑ | AB PC ↑ | AB PQ ↑ | TV Align ↑ | TA Align ↑ | ViCLIP ↑ |
| LTX-2 | 2.248 | 0.697 | 4.767 | 3.957 | 5.904 | 3.041 | 6.166 | 0.303 | 0.105 | 0.232 |
| LTX-2 + Qwen3-Omni-RL | 2.908 | 0.635 | 4.900 | 4.214 | 6.200 | 2.923 | 6.398 | 0.272 | 0.143 | 0.224 |
| LTX-2 + OmniNFT | 3.727 | 0.947 | 5.399 | 4.745 | 6.797 | 3.150 | 6.905 | 0.279 | 0.133 | 0.219 |
| LTX-2 + VA-Judger | 3.942 | 1.183 | 5.610 | 5.136 | 6.766 | 3.606 | 6.932 | 0.310 | 0.180 | 0.242 |
| Δ vs. LTX-2 | +1.693 | +0.486 | +0.843 | +1.179 | +0.862 | +0.564 | +0.766 | +0.007 | +0.075 | +0.009 |
| Δ vs. OmniNFT | +0.214 | +0.236 | +0.211 | +0.391 | -0.031 | +0.456 | +0.026 | +0.031 | +0.047 | +0.023 |
| Text Consistency | AV Consistency and Synchrony | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | TV IB ↑ | TA IB ↑ | CLIP ↑ | CLAP ↑ | AV IB ↑ | CAVP ↑ | AV Align ↑ | AVHScore ↑ | DeSync ↓ | JavisScore ↑ |
| LTX-2 | 0.355 | 0.102 | 0.303 | 0.304 | 0.091 | 0.786 | 0.192 | 0.091 | 0.430 | 0.074 |
| LTX-2 + Qwen3-Omni-RL | 0.370 | 0.140 | 0.317 | 0.354 | 0.142 | 0.799 | 0.200 | 0.142 | 0.335 | 0.122 |
| LTX-2 + OmniNFT | 0.336 | 0.134 | 0.298 | 0.394 | 0.154 | 0.800 | 0.208 | 0.146 | 0.226 | 0.122 |
| LTX-2 + VA-Judger | 0.376 | 0.179 | 0.326 | 0.425 | 0.265 | 0.799 | 0.228 | 0.261 | 0.592 | 0.230 |
| Δ vs. LTX-2 | +0.021 | +0.077 | +0.023 | +0.121 | +0.175 | +0.013 | +0.035 | +0.169 | +0.162 | +0.157 |
| Δ vs. OmniNFT | +0.040 | +0.046 | +0.028 | +0.031 | +0.111 | -0.001 | +0.020 | +0.114 | +0.366 | +0.108 |
Table 2. Green cells show the best result and underlined values show the second best.
Human preference evaluation

Human preference rates over 200 three-way comparisons. Twenty participants each evaluate all 200 prompts, yielding 4,000 selections among LTX-2, LTX-2 post-trained with OmniNFT, and LTX-2 post-trained with VA-Judger.
Generation Demos
Please turn on your audio.
Side by side comparisons of the base LTX-2 model, OmniNFT, and post training with VA-Judger. Play each clip with audio for the complete comparison.
Paper Cases
More Cases
Citation
If you find our work useful, please consider citing:
@article{huang2026vajudger,
title={VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation},
author={Huang, Yinming and Tu, Shuyuan and Yan, Xi and Yang, Zihan and Han, Jianhua and Xu, Hang and Pan, Kaihang and Jiang, Yu-Gang and Wu, Zuxuan},
journal={arXiv preprint arXiv:2608.18607},
year={2026}
}