Render to Reason: Novel-View Semantic Prediction Improves Spatial Understanding in VLMs

1University of Illinois at Urbana-Champaign, 2Meta
Render to Reason teaser figure

Render to Reason. Left: We introduce novel-view semantic rendering as an auxiliary training task, which takes input images and a target camera token, and renders the semantic layout of the unseen view. Right: On ReVSI and VSI-Bench, adding geometry features under standard QA training yields only marginal gains (+0.9, +0.1). When trained with our novel-view semantic rendering, adding geometry features yields larger gains (+2.1, +3.0), showing that the auxiliary task encourages the model to integrate geometry features more effectively.

Abstract

Recent works augment Vision-Language Models with geometry features from pretrained 3D models, expecting that the geometric signal will boost spatial reasoning. However, we find that simply fusing geometry features and training on standard spatial QA yields only marginal improvements on high-level multi-hop tasks. We attribute this gap to a training-signal problem: standard spatial QA can be largely answered from visual features and language priors, so the geometry pathway receives weak gradients and fails to integrate with the visual features. To provide a training signal that requires geometry, we propose novel-view semantic rendering as an auxiliary training task that requires the model to predict the semantic layout of an unobserved viewpoint, inspired by humans' ability to mentally simulate novel viewpoints during spatial reasoning. This task encourages joint use of both pathways: geometry provides pose-dependent visibility, while vision provides semantic content. Our auxiliary task yields consistent improvements over the geometry-augmented baseline across all three benchmarks (up to +1.6 on VSI-Bench, +2.2 on ReVSI, +2.9 on our 3D-Point-QA dataset), and our full model surpasses prior open-source methods on VSI-Bench and on ReVSI.

Model Architecture

Render2Reason model architecture

Our model jointly supports semantic novel-view rendering and standard QA, sharing a backbone and differing only in prompt and output head. Input views are encoded into visual and geometry features, fused via cross-attention, and passed to the LLM decoder. For novel-view semantic rendering, we extract a camera token from the target view via the geometry encoder and prepend it to the prompt; the LLM produces 256 output tokens, one per target patch in raster order, each mapped to a semantic class by a 2-layer MLP. For standard QA, the prompt only contains the question, and the LLM generates the answer without using the rendering MLP.

Main Results

VSI-Bench

Evaluated with 128 input views. R2R with the Qwen3-VL backbone achieves 74.4 average accuracy, the best among all compared methods, with less training data. Best results are bolded. #QA denotes the number of spatial training QAs. 9B parameters include the frozen geometry encoder.

Methods Backbone #Params #QA Avg. Numerical Answer Multiple-Choice Answer
Obj. Count Abs. Dist. Obj. Size Room Size Rel. Dist. Rel. Dir. Route Plan Appr. Order
Baseline
Chance (Frequency)–––34.062.132.029.933.125.147.928.425.2
Proprietary Models (API)
GPT-4o–––34.046.25.343.838.237.041.331.528.5
Gemini-2.5 Pro–––51.543.834.964.342.861.147.845.971.3
Open-source Fine-tuned Models
VLM-3RLLaVA-NeXT7B208K60.970.249.469.267.165.480.545.440.1
VG-LLMQwen2.5-VL9B442K62.271.456.869.069.167.983.247.432.5
SenseNova-SIInternVL38B8M68.872.053.576.872.869.680.848.576.4
GeoThinkerQwen3-VL9B1.1M72.6––––––––
Cambrian-SQwen2.57B590K67.573.250.574.972.271.176.241.880.1
Cambrian-PCambrian-S7B590K73.774.960.176.076.974.889.552.685.0
Render2ReasonInternVL3.59B394K67.372.051.971.868.268.283.545.976.4
Render2ReasonQwen3-VL9B394K74.475.064.178.375.080.088.448.585.9

ReVSI

ReVSI evaluates spatial reasoning across seven question types with manually refined annotations conditioned on the input frame count. All methods are compared with 64+ input frames. Best results are bolded. #QA denotes the number of geometry QAs generated with ground truth. 9B parameters include the geometry encoder.

Methods Backbone #Params #QA Avg. Numerical Answer Multiple-Choice Answer
Obj. Count Abs. Dist. Obj. Size Room Size Rel. Dist. Rel. Dir. Route Plan
Baseline
Chance (Frequency)–––31.452.240.117.420.925.831.930.2
Proprietary Models (64+ Frames)
GPT-5.2–––50.956.241.573.963.048.434.938.2
Gemini 3 Pro–––60.960.154.779.351.968.156.056.4
Open-source Fine-tuned Models (64+ Frames)
VSTQwen2.5-VL7B4.2M46.435.452.667.947.249.236.935.4
VLM-3RLLaVA-NeXT7B208K50.242.061.464.651.146.249.236.9
Cambrian-SQwen2.57B590K49.148.460.565.546.737.148.537.0
Cambrian-PCambrian-S7B590K52.041.468.366.745.940.248.753.1
Render2ReasonInternVL3.59B394K51.438.460.060.552.342.950.355.2
Render2ReasonQwen3-VL9B394K54.542.274.067.358.541.351.646.6

Ablation Study

We report results on VSI-Bench and ReVSI (16 frames) and on 3D-Point-QA. Sem. refers to semantic rendering. Best results are bolded. In (e), all models include VGGT features; we disable the geometry pathway at inference.

(a) Components (InternVL)

MethodsVSIReVSI3DPQA
Finetuned62.944.955.3
 + VGGT63.045.857.0
 + VGGT + Sem.64.648.059.9

(b) Components (Qwen)

MethodsVSIReVSI3DPQA
Finetuned66.050.772.2
 + VGGT67.051.577.1
 + VGGT + Sem.68.252.580.0

(c) GeoSR comparison (Qwen)

MethodsVSIReVSI3DPQA
Baseline67.051.577.1
GeoSR67.250.784.8
Ours68.252.580.0

(d) Rendering target (InternVL)

TargetVGGTVSIReVSI3DPQA
Sem.✗61.645.960.6
RGB✓62.445.260.2
RGB + Sem.✓62.646.661.1
Sem.✓64.648.059.9

(e) Geometry removal in inference

MethodsVSIReVSI3DPQA
InternVL−6.3−5.5−6.4
 + Sem.−6.7−7.3−7.0
Qwen−2.3−5.2−5.2
 + Sem.−4.2−8.9−8.7

(f) Grid resolution (InternVL)

GridVSIReVSI3DPQA
8 × 864.946.253.0
16 × 1664.648.059.9
32 × 3264.348.456.6

Effect of auxiliary task. Adding novel-view semantic rendering on top of VGGT features improves all three benchmarks on both backbones (+1.6, +2.2, and +2.9 on VSI-Bench, ReVSI, and 3D-Point-QA for InternVL; +1.2, +1.0, and +2.9 for Qwen). Without VGGT, the auxiliary task alone improves 3D-Point-QA and ReVSI but hurts VSI-Bench (−1.3), while VGGT alone yields only marginal gains, indicating the two are complementary.

Semantic vs. RGB supervision. RGB supervision improves low-level metrics, particularly point matching, but underperforms semantic supervision on high-level reasoning (VSI-Bench, ReVSI), likely because appearance-level prediction does not tie geometry to object-level semantics. We therefore use semantic rendering as our default supervision target.

Performance without geometry at inference. Removing the geometry pathway at inference hurts the auxiliary-task models more on both backbones and all three benchmarks, especially on Qwen, suggesting that our auxiliary task encourages the model to rely more on geometry features.

Comparison with GeoSR. GeoSR achieves a higher 3D-Point-QA score (84.8 vs. 80.0), but our method outperforms it on VSI-Bench (+1.0) and ReVSI (+1.8), since multi-hop spatial reasoning requires both geometric and object-level semantic understanding.

Novel-View Semantic Rendering

We evaluate auxiliary semantic rendering on 5K held-out samples where the target view occurs 1 second after the last input frame. Our model (InternVL3.5-4B) achieves 44.7% mIoU and 67.0% patch accuracy. An oracle baseline that copies the best-matching input frame's semantic map for each target frame achieves 35.6% mIoU and 57.6% patch accuracy, indicating that the model accounts for viewpoint changes rather than copying an observed layout.

Novel-view semantic rendering visualization

We visualize novel-view semantic renderings along a forward-moving camera trajectory. The first row shows the five initial RGB input frames. Each prediction uses the five RGB frames right before its target frame: the first prediction uses all five frames in row 1; the second uses frames 2–5 from row 1 and the first RGB frame in row 3; and so on.

3D-Point-QA

We construct 3D-Point-QA, a dataset with low-level reasoning tasks (with visual and textual point prompts) that should be straightforward to answer given camera poses and point clouds without object-level reasoning. The dataset includes point-to-camera distance, point-to-point distance, relative distance comparison, point matching, and 3D coordinate mapping. We extract 100K training and 5K validation QAs from ground-truth meshes in ScanNet++ following the official splits.

3D-Point-QA examples

Acknowledgement

This work is supported in part by NSF IIS grant 2312102. S.W. is supported by NSF 2331878 and 2340254, and research grants from Intel, Amazon, and IBM. This research used both the DeltaAI advanced computing and data resource, which is supported by the National Science Foundation (award OAC 2320345) and the State of Illinois, and the Delta advanced computing and data resource which is supported by the National Science Foundation (award OAC 2005572) and the State of Illinois. Delta and DeltaAI are joint efforts of the University of Illinois Urbana-Champaign and its National Center for Supercomputing Applications.