<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Research on Infinite Script</title><link>https://www.infinitescript.com/categories/research/</link><description>Recent content in Research on Infinite Script</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Tue, 09 Dec 2025 00:00:00 +0000</lastBuildDate><atom:link href="https://www.infinitescript.com/categories/research/index.xml" rel="self" type="application/rss+xml"/><item><title>DynamicVLA</title><link>https://www.infinitescript.com/project/dynamic-vla/</link><pubDate>Tue, 09 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.infinitescript.com/project/dynamic-vla/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; DynamicVLA enables open-ended dynamic object manipulation by pairing a compact 0.4B VLM with low-latency Continuous Inference and Latent-aware Action Streaming, evaluated at scale through the new DOM benchmark in both simulation and the real world.&lt;/p&gt;&#10;&lt;div class="gallery gallery-slides"&gt;&#10; &lt;div class="overlay"&gt;&#10; &lt;img src="https://www.infinitescript.com/images/loading.gif" alt=""&gt;&#10; &lt;/div&gt; &#10; &lt;div class="gallery-inner"&gt;&#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/Franka-Orange.webp" data-alt="DynamicVLA (1/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/Franka-Orange.webp" alt="DynamicVLA (1/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Apple-Green.webp" data-alt="DynamicVLA (2/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Apple-Green.webp" alt="DynamicVLA (2/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/PiPER-PingPong.webp" data-alt="DynamicVLA (3/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/PiPER-PingPong.webp" alt="DynamicVLA (3/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/PiPER-FoodCan.webp" data-alt="DynamicVLA (4/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/PiPER-FoodCan.webp" alt="DynamicVLA (4/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Cola.webp" data-alt="DynamicVLA (5/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Cola.webp" alt="DynamicVLA (5/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Orange.webp" data-alt="DynamicVLA (6/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Orange.webp" alt="DynamicVLA (6/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Apple-Red.webp" data-alt="DynamicVLA (7/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Apple-Red.webp" alt="DynamicVLA (7/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/Franka-Apple-Red.webp" data-alt="DynamicVLA (8/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/Franka-Apple-Red.webp" alt="DynamicVLA (8/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/Franka-Potato.webp" data-alt="DynamicVLA (9/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/Franka-Potato.webp" alt="DynamicVLA (9/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/Franka-Jujube.webp" data-alt="DynamicVLA (10/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/Franka-Jujube.webp" alt="DynamicVLA (10/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/Franka-Lemon.webp" data-alt="DynamicVLA (11/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/Franka-Lemon.webp" alt="DynamicVLA (11/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/Franka-Avocado.webp" data-alt="DynamicVLA (12/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/Franka-Avocado.webp" alt="DynamicVLA (12/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &#10; &lt;/div&gt; &#10; &#10; &lt;button class="carousel-control carousel-control-prev" type="button" tabindex="0" aria-label="Previous image"&gt;&#10; &lt;span class="carousel-control-prev-icon" aria-hidden="true"&gt;&lt;/span&gt;&#10; &lt;/button&gt;&#10; &lt;button class="carousel-control carousel-control-next" type="button" tabindex="0" aria-label="Next image"&gt;&#10; &lt;span class="carousel-control-next-icon" aria-hidden="true"&gt;&lt;/span&gt;&#10; &lt;/button&gt;&#10; &#10; &lt;span class="gallery-counter" aria-hidden="true"&gt;&lt;b&gt;1&lt;/b&gt;&amp;thinsp;/&amp;thinsp;12&lt;/span&gt;&#10; &#10;&lt;/div&gt; &#10;&#10;&lt;h2 id="highlights"&gt;Highlights&lt;/h2&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/DynamicVLA/DynamicVLA-Teaser.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/DynamicVLA/DynamicVLA-Teaser.webp" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;</description></item><item><title>CityDreamer4D</title><link>https://www.infinitescript.com/project/city-dreamer-4d/</link><pubDate>Tue, 31 Dec 2024 22:00:00 +0000</pubDate><guid>https://www.infinitescript.com/project/city-dreamer-4d/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; CityDreamer4D is a framework for unbounded 4D city generation that decouples static and dynamic scenes, achieving superior realism, multi-view consistency, and diverse styles.&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;3D scene generation has garnered growing attention in recent years and has made significant progress.&#10;Generating 4D cities is more challenging than 3D scenes due to the presence of structurally complex, visually diverse objects like buildings and vehicles, and heightened human sensitivity to distortions in urban environments.&#10;To tackle these issues, we propose &lt;strong&gt;CityDreamer4D&lt;/strong&gt;, a compositional generative model specifically tailored for generating unbounded 4D cities. Our main insights are&#10;&lt;strong&gt;1)&lt;/strong&gt; 4D city generation should separate dynamic objects (&lt;em&gt;e.g.&lt;/em&gt;, vehicles) from static scenes (&lt;em&gt;e.g.&lt;/em&gt;, buildings and roads), and&#10;&lt;strong&gt;2)&lt;/strong&gt; all objects in the 4D scene should be composed of different types of neural fields for buildings, vehicles, and background stuff.&#10;Specifically, we propose Traffic Scenario Generator and Unbounded Layout Generator to produce dynamic traffic scenarios and static city layouts using a highly compact BEV representation.&#10;Objects in 4D cities are generated by combining stuff-oriented and instance-oriented neural fields for background stuff, buildings, and vehicles.&#10;To suit the distinct characteristics of background stuff and instances, the neural fields employ customized generative hash grids and periodic positional embeddings as scene parameterizations.&#10;Furthermore, we offer a comprehensive suite of datasets for city generation, including OSM, GoogleEarth, and CityTopia.&#10;The OSM dataset provides a variety of real-world city layouts, while the Google Earth and CityTopia datasets deliver large-scale, high-quality city imagery complete with 3D instance annotations.&#10;Leveraging its compositional design, CityDreamer4D supports a range of downstream applications, such as instance editing, city stylization, and urban simulation, while delivering state-of-the-art performance in generating realistic 4D cities.&lt;/p&gt;</description></item><item><title>GaussianCity</title><link>https://www.infinitescript.com/project/gaussian-city/</link><pubDate>Fri, 24 May 2024 08:00:00 +0000</pubDate><guid>https://www.infinitescript.com/project/gaussian-city/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; GaussianCity is a framework for efficient unbounded 3D city generation using 3D Gaussian Splatting.&#10;&lt;br&gt;&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/GaussianCity/GaussianCity-Teaser.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/GaussianCity/GaussianCity-Teaser.webp" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;3D city generation with NeRF-based methods shows promising generation results but is computationally inefficient. Recently 3D Gaussian Splatting (3D-GS) has emerged as a highly efficient alternative for object-level 3D generation. However, adapting 3D-GS from finite-scale 3D objects and humans to infinite-scale 3D cities is non-trivial. Unbounded 3D city generation entails significant storage overhead (out-of-memory issues), arising from the need to expand points to billions, often demanding hundreds of Gigabytes of VRAM for a city scene spanning 10km&lt;sup&gt;2&lt;/sup&gt;. In this paper, we propose &lt;strong&gt;GaussianCity&lt;/strong&gt;, a generative Gaussian Splatting framework dedicated to efficiently synthesize unbounded 3D cities with a single feed-forward pass. Our key insights are two-fold: &lt;strong&gt;1)&lt;/strong&gt; Compact 3D Scene Representation: We introduce BEV-Point as a highly compact intermediate representation, ensuring that the growth in VRAM usage for unbounded scenes remains constant, thus enabling unbounded city generation. &lt;strong&gt;2)&lt;/strong&gt; Spatial-aware Gaussian Attribute Decoder: We present spatial-aware BEV-Point decoder to produce 3D Gaussian attributes, which leverages Point Serializer to integrate the structural and contextual characteristics of BEV points. Extensive experiments demonstrate that GaussianCity achieves state-of-the-art results in both drone-view and street-view 3D city generation. Notably, compared to CityDreamer, GaussianCity exhibits superior performance with a speedup of 60 times (10.72 FPS v.s. 0.18 FPS).&lt;/p&gt;</description></item><item><title>CityDreamer</title><link>https://www.infinitescript.com/project/city-dreamer/</link><pubDate>Fri, 01 Sep 2023 08:00:00 +0000</pubDate><guid>https://www.infinitescript.com/project/city-dreamer/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; CityDreamer learns to generate unbounded 3D cities from Google Earth imagery and OpenStreetMap.&#10;&lt;br&gt;&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/CityDreamer/CityDreamer-Teaser.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/CityDreamer/CityDreamer-Teaser.webp" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;In recent years, extensive research has focused on 3D natural scene generation, but the domain of 3D city generation has not received as much exploration. This is due to the greater challenges posed by 3D city generation, mainly because humans are more sensitive to structural distortions in urban environments. Additionally, generating 3D cities is more complex than 3D natural scenes since buildings, as objects of the same class, exhibit a wider range of appearances compared to the relatively consistent appearance of objects like trees in natural scenes. To address these challenges, we propose CityDreamer, a compositional generative model designed specifically for unbounded 3D cities, which separates the generation of building instances from other background objects, such as roads, green lands, and water areas, into distinct modules. Furthermore, we construct two datasets, OSM and GoogleEarth, containing a vast amount of real-world city imagery to enhance the realism of the generated 3D cities both in their layout and appearance. Through extensive experiments, CityDreamer has proven its superiority over state-of-the-art methods in generating a wide range of lifelike 3D cities.&lt;/p&gt;</description></item><item><title>RMNet</title><link>https://www.infinitescript.com/project/rmnet/</link><pubDate>Sat, 06 Mar 2021 16:32:00 +0000</pubDate><guid>https://www.infinitescript.com/project/rmnet/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; RMNet segments video objects by local-to-local matching, memorizing only the regions where the target appeared in past frames and tracking the query region with optical flow, which reduces both similar-object mismatching and computational cost.&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/RMNet/RMNet-Overview.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/RMNet/RMNet-Overview.webp" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;Recently, several Space-Time Memory based networks have shown that the object cues (e.g. video frames as well as the segmented object masks) from the past frames are useful for segmenting objects in the current frame. However, these methods exploit the information from the memory by global-to-global matching between the current and past frames, which lead to mismatching to similar objects and high computational complexity. To address these problems, we propose a novel local-to-local matching solution for semi-supervised VOS, namely Regional Memory Network (RMNet). In RMNet, the precise regional memory is constructed by memorizing local regions where the target objects appear in the past frames. For the current query frame, the query regions are tracked and predicted based on the optical flow estimated from the previous frame. The proposed local-to-local matching effectively alleviates the ambiguity of similar objects in both memory and query frames, which allows the information to be passed from the regional memory to the query region efficiently and effectively. Experimental results indicate that the proposed RMNet performs favorably against state-of-the-art methods on the DAVIS and YouTube-VOS datasets.&lt;/p&gt;</description></item><item><title>GRNet</title><link>https://www.infinitescript.com/project/grnet/</link><pubDate>Fri, 03 Jul 2020 06:42:00 +0000</pubDate><guid>https://www.infinitescript.com/project/grnet/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; GRNet completes dense 3D point clouds by regularizing them into 3D grids, with differentiable Gridding, Gridding Reverse, and Cubic Feature Sampling layers and a Gridding Loss that recovers fine details on ShapeNet, Completion3D, and KITTI.&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/GRNet/GRNet-Overview.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/GRNet/GRNet-Overview.webp" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;Estimating the complete 3D point cloud from an incomplete one is a key problem in many vision and robotics applications. Mainstream methods (e.g., PCN and TopNet) use Multi-layer Perceptrons (MLPs) to directly process point clouds, which may cause the loss of details because the structural and context of point clouds are not fully considered. To solve this problem, we introduce 3D grids as intermediate representations to regularize unordered point clouds. We therefore propose a novel Gridding Residual Network (GRNet) for point cloud completion. In particular, we devise two novel differentiable layers, named Gridding and Gridding Reverse, to convert between point clouds and 3D grids without losing structural information. We also present the differentiable Cubic Feature Sampling layer to extract features of neighboring points, which preserves context information. In addition, we design a new loss function, namely Gridding Loss, to calculate the L1 distance between the 3D grids of the predicted and ground truth point clouds, which is helpful to recover details. Experimental results indicate that the proposed GRNet performs favorably against state-of-the-art methods on the ShapeNet, Completion3D, and KITTI benchmarks.&lt;/p&gt;</description></item><item><title>Stereo 3D Reconstruction</title><link>https://www.infinitescript.com/project/stereo-3d-reconstruction/</link><pubDate>Sun, 27 Oct 2019 06:42:00 +0000</pubDate><guid>https://www.infinitescript.com/project/stereo-3d-reconstruction/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; This work reconstructs an object&amp;rsquo;s 3D shape from a pair of stereo images by reasoning about bidirectional disparities and cross-view feature correspondences, and introduces StereoShapeNet, a benchmark of 1,052,976 stereo pairs rendered from ShapeNet.&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/Stereo-3D-Reconstruction/Stereo-3D-Reconstruction-Overview.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/Stereo-3D-Reconstruction/Stereo-3D-Reconstruction-Overview.webp" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;Inferring the complete 3D shape of an object from an RGB image has shown impressive results, however, existing methods rely primarily on recognizing the most similar 3D model from the training set to solve the problem. These methods suffer from poor generalization and may lead to low-quality reconstructions for unseen objects. Nowadays, stereo cameras are pervasive in emerging devices such as dual-lens smartphones and robots, which enables the use of the two-view nature of stereo images to explore the 3D structure and thus improve the reconstruction performance. In this paper, we propose a new deep learning framework for reconstructing the 3D shape of an object from a pair of stereo images, which reasons about the 3D structure of the object by taking bidirectional disparities and feature correspondences between the two views into account. Besides, we present a large-scale synthetic benchmarking dataset, namely StereoShapeNet, containing 1,052,976 pairs of stereo images rendered from ShapeNet along with the corresponding bidirectional depth and disparity maps. Experimental results on the StereoShapeNet benchmark demonstrate that the proposed framework outperforms the state-of-the-art methods.&lt;/p&gt;</description></item><item><title>Pix2Vox</title><link>https://www.infinitescript.com/project/pix2vox/</link><pubDate>Tue, 30 Apr 2019 14:58:00 +0000</pubDate><guid>https://www.infinitescript.com/project/pix2vox/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Pix2Vox reconstructs an object&amp;rsquo;s 3D shape from single-view or multi-view images, using a context-aware fusion module that selects the best-reconstructed parts across views to produce order-invariant results 24 times faster than 3D-R2N2.&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/Pix2Vox/Pix2Vox-Overview.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/Pix2Vox/Pix2Vox-Overview.webp" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;Recovering the 3D representation of an object from single-view or multi-view RGB images by deep neural networks has attracted increasing attention in the past few years. Several mainstream works (e.g., 3D-R2N2) use recurrent neural networks (RNNs) to fuse multiple feature maps extracted from input images sequentially. However, when given the same set of input images with different orders, RNN-based approaches are unable to produce consistent reconstruction results. Moreover, due to long-term memory loss, RNNs cannot fully exploit input images to refine reconstruction results. To solve these problems, we propose a novel framework for single-view and multi-view 3D reconstruction, named Pix2Vox. By using a well-designed encoder-decoder, it generates a coarse 3D volume from each input image. Then, a context-aware fusion module is introduced to adaptively select high-quality reconstructions for each part (e.g., table legs) from different coarse 3D volumes to obtain a fused 3D volume. Finally, a refiner further refines the fused 3D volume to generate the final output. Experimental results on the ShapeNet and Pix3D benchmarks indicate that the proposed Pix2Vox outperforms state-of-the-arts by a large margin. Furthermore, the proposed method is 24 times faster than 3D-R2N2 in terms of backward inference time. The experiments on ShapeNet unseen 3D categories have shown the superior generalization abilities of our method.&lt;/p&gt;</description></item><item><title>Weighted Voxel</title><link>https://www.infinitescript.com/project/weighted-voxel/</link><pubDate>Sat, 03 Feb 2018 03:54:00 +0000</pubDate><guid>https://www.infinitescript.com/project/weighted-voxel/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Weighted Voxel replaces the zero-one occupancy grid with a richer voxel representation that retains structural information, improving reconstruction quality while taking less time to train.&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/Weighted-Voxel/Weighted-Voxel-Overview.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/Weighted-Voxel/Weighted-Voxel-Overview.webp" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;3D reconstruction has been attracting increasing attention in the past few years. With the surge of deep neural networks, the performance of 3D reconstruction has been improved significantly. However, the voxel reconstructed by extant approaches usually contains lots of noise and leads to heavy computation. In this paper, we define a new voxel representation, named Weighted Voxel. It provides more abundant information, facilitating the subsequent learning and generalization steps. Unlike regular voxel which consists of zero-one, the proposed Weighted Voxel makes full use of the structure information of voxels. Experimental results demonstrate that Weighted Voxel not only performs better in reconstruction but also takes less time in training.&lt;/p&gt;</description></item></channel></rss>