<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Projects on Infinite Script</title><link>https://www.infinitescript.com/projects/</link><description>Recent content in Projects on Infinite Script</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Tue, 09 Dec 2025 00:00:00 +0000</lastBuildDate><atom:link href="https://www.infinitescript.com/projects/index.xml" rel="self" type="application/rss+xml"/><item><title>DynamicVLA</title><link>https://www.infinitescript.com/project/dynamic-vla/</link><pubDate>Tue, 09 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.infinitescript.com/project/dynamic-vla/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; DynamicVLA enables open-ended dynamic object manipulation by pairing a compact 0.4B VLM with low-latency Continuous Inference and Latent-aware Action Streaming, evaluated at scale through the new DOM benchmark in both simulation and the real world.&lt;/p&gt;&#10;&lt;div class="gallery gallery-slides"&gt;&#10; &lt;div class="overlay"&gt;&#10; &lt;img src="https://www.infinitescript.com/images/loading.gif" width="188" height="88" alt=""&gt;&#10; &lt;/div&gt; &#10; &lt;div class="gallery-inner"&gt;&#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/Franka-Orange.webp" data-alt="DynamicVLA (1/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/Franka-Orange.webp" width="360" height="203" alt="DynamicVLA (1/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Apple-Green.webp" data-alt="DynamicVLA (2/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Apple-Green.webp" width="360" height="203" alt="DynamicVLA (2/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/PiPER-PingPong.webp" data-alt="DynamicVLA (3/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/PiPER-PingPong.webp" width="360" height="203" alt="DynamicVLA (3/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/PiPER-FoodCan.webp" data-alt="DynamicVLA (4/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/PiPER-FoodCan.webp" width="360" height="203" alt="DynamicVLA (4/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Cola.webp" data-alt="DynamicVLA (5/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Cola.webp" width="360" height="203" alt="DynamicVLA (5/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Orange.webp" data-alt="DynamicVLA (6/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Orange.webp" width="360" height="203" alt="DynamicVLA (6/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Apple-Red.webp" data-alt="DynamicVLA (7/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/PiPER-Apple-Red.webp" width="360" height="203" alt="DynamicVLA (7/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/Franka-Apple-Red.webp" data-alt="DynamicVLA (8/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/Franka-Apple-Red.webp" width="360" height="203" alt="DynamicVLA (8/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/Franka-Potato.webp" data-alt="DynamicVLA (9/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/Franka-Potato.webp" width="360" height="203" alt="DynamicVLA (9/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/Franka-Jujube.webp" data-alt="DynamicVLA (10/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/Franka-Jujube.webp" width="360" height="203" alt="DynamicVLA (10/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/Franka-Lemon.webp" data-alt="DynamicVLA (11/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/Franka-Lemon.webp" width="360" height="203" alt="DynamicVLA (11/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &lt;div class="gallery-item col-12 col-md-6 col-lg-4" url="https://www.infinitescript.com/projects/DynamicVLA/Franka-Avocado.webp" data-alt="DynamicVLA (12/12)"&gt;&#10; &lt;noscript&gt;&lt;img src="https://www.infinitescript.com/projects/DynamicVLA/Franka-Avocado.webp" width="360" height="203" alt="DynamicVLA (12/12)" loading="lazy"&gt;&lt;/noscript&gt;&#10; &#10; &lt;/div&gt; &#10; &#10; &lt;/div&gt; &#10; &#10; &lt;button class="carousel-control carousel-control-prev" type="button" tabindex="0" aria-label="Previous image"&gt;&#10; &lt;span class="carousel-control-prev-icon" aria-hidden="true"&gt;&lt;/span&gt;&#10; &lt;/button&gt;&#10; &lt;button class="carousel-control carousel-control-next" type="button" tabindex="0" aria-label="Next image"&gt;&#10; &lt;span class="carousel-control-next-icon" aria-hidden="true"&gt;&lt;/span&gt;&#10; &lt;/button&gt;&#10; &#10; &lt;span class="gallery-counter" aria-hidden="true"&gt;&lt;b&gt;1&lt;/b&gt;&amp;thinsp;/&amp;thinsp;12&lt;/span&gt;&#10; &#10;&lt;/div&gt; &#10;&#10;&lt;h2 id="highlights"&gt;Highlights&lt;/h2&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/DynamicVLA/DynamicVLA-Teaser.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/DynamicVLA/DynamicVLA-Teaser.webp" width="1920" height="821" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;</description></item><item><title>CityDreamer4D</title><link>https://www.infinitescript.com/project/city-dreamer-4d/</link><pubDate>Tue, 31 Dec 2024 22:00:00 +0000</pubDate><guid>https://www.infinitescript.com/project/city-dreamer-4d/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; CityDreamer4D is a framework for unbounded 4D city generation that decouples static and dynamic scenes, achieving superior realism, multi-view consistency, and diverse styles.&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;3D scene generation has garnered growing attention in recent years and has made significant progress.&#10;Generating 4D cities is more challenging than 3D scenes due to the presence of structurally complex, visually diverse objects like buildings and vehicles, and heightened human sensitivity to distortions in urban environments.&#10;To tackle these issues, we propose &lt;strong&gt;CityDreamer4D&lt;/strong&gt;, a compositional generative model specifically tailored for generating unbounded 4D cities. Our main insights are&#10;&lt;strong&gt;1)&lt;/strong&gt; 4D city generation should separate dynamic objects (&lt;em&gt;e.g.&lt;/em&gt;, vehicles) from static scenes (&lt;em&gt;e.g.&lt;/em&gt;, buildings and roads), and&#10;&lt;strong&gt;2)&lt;/strong&gt; all objects in the 4D scene should be composed of different types of neural fields for buildings, vehicles, and background stuff.&#10;Specifically, we propose Traffic Scenario Generator and Unbounded Layout Generator to produce dynamic traffic scenarios and static city layouts using a highly compact BEV representation.&#10;Objects in 4D cities are generated by combining stuff-oriented and instance-oriented neural fields for background stuff, buildings, and vehicles.&#10;To suit the distinct characteristics of background stuff and instances, the neural fields employ customized generative hash grids and periodic positional embeddings as scene parameterizations.&#10;Furthermore, we offer a comprehensive suite of datasets for city generation, including OSM, GoogleEarth, and CityTopia.&#10;The OSM dataset provides a variety of real-world city layouts, while the Google Earth and CityTopia datasets deliver large-scale, high-quality city imagery complete with 3D instance annotations.&#10;Leveraging its compositional design, CityDreamer4D supports a range of downstream applications, such as instance editing, city stylization, and urban simulation, while delivering state-of-the-art performance in generating realistic 4D cities.&lt;/p&gt;</description></item><item><title>GaussianCity</title><link>https://www.infinitescript.com/project/gaussian-city/</link><pubDate>Fri, 24 May 2024 08:00:00 +0000</pubDate><guid>https://www.infinitescript.com/project/gaussian-city/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; GaussianCity is a framework for efficient unbounded 3D city generation using 3D Gaussian Splatting.&#10;&lt;br&gt;&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/GaussianCity/GaussianCity-Teaser.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/GaussianCity/GaussianCity-Teaser.webp" width="1920" height="640" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;3D city generation with NeRF-based methods shows promising generation results but is computationally inefficient. Recently 3D Gaussian Splatting (3D-GS) has emerged as a highly efficient alternative for object-level 3D generation. However, adapting 3D-GS from finite-scale 3D objects and humans to infinite-scale 3D cities is non-trivial. Unbounded 3D city generation entails significant storage overhead (out-of-memory issues), arising from the need to expand points to billions, often demanding hundreds of Gigabytes of VRAM for a city scene spanning 10km&lt;sup&gt;2&lt;/sup&gt;. In this paper, we propose &lt;strong&gt;GaussianCity&lt;/strong&gt;, a generative Gaussian Splatting framework dedicated to efficiently synthesize unbounded 3D cities with a single feed-forward pass. Our key insights are two-fold: &lt;strong&gt;1)&lt;/strong&gt; Compact 3D Scene Representation: We introduce BEV-Point as a highly compact intermediate representation, ensuring that the growth in VRAM usage for unbounded scenes remains constant, thus enabling unbounded city generation. &lt;strong&gt;2)&lt;/strong&gt; Spatial-aware Gaussian Attribute Decoder: We present spatial-aware BEV-Point decoder to produce 3D Gaussian attributes, which leverages Point Serializer to integrate the structural and contextual characteristics of BEV points. Extensive experiments demonstrate that GaussianCity achieves state-of-the-art results in both drone-view and street-view 3D city generation. Notably, compared to CityDreamer, GaussianCity exhibits superior performance with a speedup of 60 times (10.72 FPS v.s. 0.18 FPS).&lt;/p&gt;</description></item><item><title>Infinite Servers</title><link>https://www.infinitescript.com/project/infinite-servers/</link><pubDate>Fri, 02 Feb 2024 11:41:27 +0000</pubDate><guid>https://www.infinitescript.com/project/infinite-servers/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;&#10;&lt;p&gt;&lt;strong&gt;Infinite Servers&lt;/strong&gt; is a self-hosted server fleet monitor built around a single design constraint: the agent must run completely unprivileged. Each monitored host installs a PHP script that runs as &lt;code&gt;nobody&lt;/code&gt;, reads &lt;code&gt;/proc&lt;/code&gt;, and pushes status every 15 seconds to a central dashboard. No root access, no persistent daemon with elevated privileges, no external database server.&lt;/p&gt;&#10;&lt;p&gt;This design was motivated by the security risks that plague popular alternatives. Tools like &lt;a href="https://github.com/naiba/nezha"&gt;Nezha&lt;/a&gt; require agents to run as root, which creates a severe blast-radius problem: if the dashboard is compromised, an attacker can push arbitrary commands to every connected host with full system privileges. Two recent CVEs illustrate just how real this risk is:&lt;/p&gt;</description></item><item><title>CityDreamer</title><link>https://www.infinitescript.com/project/city-dreamer/</link><pubDate>Fri, 01 Sep 2023 08:00:00 +0000</pubDate><guid>https://www.infinitescript.com/project/city-dreamer/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; CityDreamer learns to generate unbounded 3D cities from Google Earth imagery and OpenStreetMap.&#10;&lt;br&gt;&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/CityDreamer/CityDreamer-Teaser.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/CityDreamer/CityDreamer-Teaser.webp" width="1280" height="420" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;In recent years, extensive research has focused on 3D natural scene generation, but the domain of 3D city generation has not received as much exploration. This is due to the greater challenges posed by 3D city generation, mainly because humans are more sensitive to structural distortions in urban environments. Additionally, generating 3D cities is more complex than 3D natural scenes since buildings, as objects of the same class, exhibit a wider range of appearances compared to the relatively consistent appearance of objects like trees in natural scenes. To address these challenges, we propose CityDreamer, a compositional generative model designed specifically for unbounded 3D cities, which separates the generation of building instances from other background objects, such as roads, green lands, and water areas, into distinct modules. Furthermore, we construct two datasets, OSM and GoogleEarth, containing a vast amount of real-world city imagery to enhance the realism of the generated 3D cities both in their layout and appearance. Through extensive experiments, CityDreamer has proven its superiority over state-of-the-art methods in generating a wide range of lifelike 3D cities.&lt;/p&gt;</description></item><item><title>RMNet</title><link>https://www.infinitescript.com/project/rmnet/</link><pubDate>Sat, 06 Mar 2021 16:32:00 +0000</pubDate><guid>https://www.infinitescript.com/project/rmnet/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; RMNet segments video objects by local-to-local matching, memorizing only the regions where the target appeared in past frames and tracking the query region with optical flow, which reduces both similar-object mismatching and computational cost.&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/RMNet/RMNet-Overview.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/RMNet/RMNet-Overview.webp" width="1920" height="925" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;Recently, several Space-Time Memory based networks have shown that the object cues (e.g. video frames as well as the segmented object masks) from the past frames are useful for segmenting objects in the current frame. However, these methods exploit the information from the memory by global-to-global matching between the current and past frames, which lead to mismatching to similar objects and high computational complexity. To address these problems, we propose a novel local-to-local matching solution for semi-supervised VOS, namely Regional Memory Network (RMNet). In RMNet, the precise regional memory is constructed by memorizing local regions where the target objects appear in the past frames. For the current query frame, the query regions are tracked and predicted based on the optical flow estimated from the previous frame. The proposed local-to-local matching effectively alleviates the ambiguity of similar objects in both memory and query frames, which allows the information to be passed from the regional memory to the query region efficiently and effectively. Experimental results indicate that the proposed RMNet performs favorably against state-of-the-art methods on the DAVIS and YouTube-VOS datasets.&lt;/p&gt;</description></item><item><title>GRNet</title><link>https://www.infinitescript.com/project/grnet/</link><pubDate>Fri, 03 Jul 2020 06:42:00 +0000</pubDate><guid>https://www.infinitescript.com/project/grnet/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; GRNet completes dense 3D point clouds by regularizing them into 3D grids, with differentiable Gridding, Gridding Reverse, and Cubic Feature Sampling layers and a Gridding Loss that recovers fine details on ShapeNet, Completion3D, and KITTI.&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/GRNet/GRNet-Overview.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/GRNet/GRNet-Overview.webp" width="2448" height="915" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;Estimating the complete 3D point cloud from an incomplete one is a key problem in many vision and robotics applications. Mainstream methods (e.g., PCN and TopNet) use Multi-layer Perceptrons (MLPs) to directly process point clouds, which may cause the loss of details because the structural and context of point clouds are not fully considered. To solve this problem, we introduce 3D grids as intermediate representations to regularize unordered point clouds. We therefore propose a novel Gridding Residual Network (GRNet) for point cloud completion. In particular, we devise two novel differentiable layers, named Gridding and Gridding Reverse, to convert between point clouds and 3D grids without losing structural information. We also present the differentiable Cubic Feature Sampling layer to extract features of neighboring points, which preserves context information. In addition, we design a new loss function, namely Gridding Loss, to calculate the L1 distance between the 3D grids of the predicted and ground truth point clouds, which is helpful to recover details. Experimental results indicate that the proposed GRNet performs favorably against state-of-the-art methods on the ShapeNet, Completion3D, and KITTI benchmarks.&lt;/p&gt;</description></item><item><title>DataNinja</title><link>https://www.infinitescript.com/project/dataninja/</link><pubDate>Tue, 18 Feb 2020 23:23:00 +0000</pubDate><guid>https://www.infinitescript.com/project/dataninja/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;&#10;&lt;p&gt;&lt;strong&gt;DataNinja&lt;/strong&gt; is a full-stack machine learning workbench designed for object detection tasks. It provides a browser-based interface to manage datasets, configure and train detection models, evaluate results, and, most distinctively, push trained models to embedded hardware in a single workflow.&lt;/p&gt;&#10;&lt;p&gt;The motivation behind the project was to bridge the gap between the server side (where GPUs live and models are trained) and the edge side (where inference actually runs, often on low-power embedded SoCs or FPGAs). Most open-source ML tools at the time handled one or the other; DataNinja treats them as a single, connected pipeline.&lt;/p&gt;</description></item><item><title>Stereo 3D Reconstruction</title><link>https://www.infinitescript.com/project/stereo-3d-reconstruction/</link><pubDate>Sun, 27 Oct 2019 06:42:00 +0000</pubDate><guid>https://www.infinitescript.com/project/stereo-3d-reconstruction/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; This work reconstructs an object&amp;rsquo;s 3D shape from a pair of stereo images by reasoning about bidirectional disparities and cross-view feature correspondences, and introduces StereoShapeNet, a benchmark of 1,052,976 stereo pairs rendered from ShapeNet.&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/Stereo-3D-Reconstruction/Stereo-3D-Reconstruction-Overview.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/Stereo-3D-Reconstruction/Stereo-3D-Reconstruction-Overview.webp" width="2402" height="1074" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;Inferring the complete 3D shape of an object from an RGB image has shown impressive results, however, existing methods rely primarily on recognizing the most similar 3D model from the training set to solve the problem. These methods suffer from poor generalization and may lead to low-quality reconstructions for unseen objects. Nowadays, stereo cameras are pervasive in emerging devices such as dual-lens smartphones and robots, which enables the use of the two-view nature of stereo images to explore the 3D structure and thus improve the reconstruction performance. In this paper, we propose a new deep learning framework for reconstructing the 3D shape of an object from a pair of stereo images, which reasons about the 3D structure of the object by taking bidirectional disparities and feature correspondences between the two views into account. Besides, we present a large-scale synthetic benchmarking dataset, namely StereoShapeNet, containing 1,052,976 pairs of stereo images rendered from ShapeNet along with the corresponding bidirectional depth and disparity maps. Experimental results on the StereoShapeNet benchmark demonstrate that the proposed framework outperforms the state-of-the-art methods.&lt;/p&gt;</description></item><item><title>Pix2Vox</title><link>https://www.infinitescript.com/project/pix2vox/</link><pubDate>Tue, 30 Apr 2019 14:58:00 +0000</pubDate><guid>https://www.infinitescript.com/project/pix2vox/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Pix2Vox reconstructs an object&amp;rsquo;s 3D shape from single-view or multi-view images, using a context-aware fusion module that selects the best-reconstructed parts across views to produce order-invariant results 24 times faster than 3D-R2N2.&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/Pix2Vox/Pix2Vox-Overview.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/Pix2Vox/Pix2Vox-Overview.webp" width="1800" height="520" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;Recovering the 3D representation of an object from single-view or multi-view RGB images by deep neural networks has attracted increasing attention in the past few years. Several mainstream works (e.g., 3D-R2N2) use recurrent neural networks (RNNs) to fuse multiple feature maps extracted from input images sequentially. However, when given the same set of input images with different orders, RNN-based approaches are unable to produce consistent reconstruction results. Moreover, due to long-term memory loss, RNNs cannot fully exploit input images to refine reconstruction results. To solve these problems, we propose a novel framework for single-view and multi-view 3D reconstruction, named Pix2Vox. By using a well-designed encoder-decoder, it generates a coarse 3D volume from each input image. Then, a context-aware fusion module is introduced to adaptively select high-quality reconstructions for each part (e.g., table legs) from different coarse 3D volumes to obtain a fused 3D volume. Finally, a refiner further refines the fused 3D volume to generate the final output. Experimental results on the ShapeNet and Pix3D benchmarks indicate that the proposed Pix2Vox outperforms state-of-the-arts by a large margin. Furthermore, the proposed method is 24 times faster than 3D-R2N2 in terms of backward inference time. The experiments on ShapeNet unseen 3D categories have shown the superior generalization abilities of our method.&lt;/p&gt;</description></item><item><title>Weighted Voxel</title><link>https://www.infinitescript.com/project/weighted-voxel/</link><pubDate>Sat, 03 Feb 2018 03:54:00 +0000</pubDate><guid>https://www.infinitescript.com/project/weighted-voxel/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Weighted Voxel replaces the zero-one occupancy grid with a richer voxel representation that retains structural information, improving reconstruction quality while taking less time to train.&lt;/p&gt;&#10;&lt;p&gt;&#10;&#10;&lt;a href="https://www.infinitescript.com/projects/Weighted-Voxel/Weighted-Voxel-Overview.webp" data-fancybox data-caption="Teaser"&gt;&#10; &lt;img src="https://www.infinitescript.com/projects/Weighted-Voxel/Weighted-Voxel-Overview.webp" width="1500" height="600" alt="Teaser" loading="lazy"&gt;&#10;&lt;/a&gt;&#10;&#10;&#10;&lt;/p&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;3D reconstruction has been attracting increasing attention in the past few years. With the surge of deep neural networks, the performance of 3D reconstruction has been improved significantly. However, the voxel reconstructed by extant approaches usually contains lots of noise and leads to heavy computation. In this paper, we define a new voxel representation, named Weighted Voxel. It provides more abundant information, facilitating the subsequent learning and generalization steps. Unlike regular voxel which consists of zero-one, the proposed Weighted Voxel makes full use of the structure information of voxels. Experimental results demonstrate that Weighted Voxel not only performs better in reconstruction but also takes less time in training.&lt;/p&gt;</description></item><item><title>Similar Patient Finder</title><link>https://www.infinitescript.com/project/similar-patient-finder/</link><pubDate>Sun, 18 Jun 2017 09:23:00 +0000</pubDate><guid>https://www.infinitescript.com/project/similar-patient-finder/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;&#10;&lt;p&gt;&lt;strong&gt;Similar Patient Finder (SPF)&lt;/strong&gt; is a web-based clinical decision support tool that helps physicians explore patient similarities at the molecular level. Given a dataset of high-throughput gene expression profiles, optionally paired with clinical metadata, SPF runs a configurable analysis pipeline and renders the results as an interactive 2D patient map, making it easier to identify disease subtypes, staging, and appropriate therapy options.&lt;/p&gt;&#10;&lt;p&gt;The motivation was that raw genomic data is too high-dimensional for direct inspection. SPF bridges the gap between bioinformatics methods and clinical workflows by wrapping a full preprocessing-to-visualization pipeline in a browser-based interface that requires no command-line work from the physician.&lt;/p&gt;</description></item><item><title>Verwandlung Online Judge</title><link>https://www.infinitescript.com/project/verwandlung-online-judge/</link><pubDate>Thu, 30 Apr 2015 18:45:00 +0000</pubDate><guid>https://www.infinitescript.com/project/verwandlung-online-judge/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;&#10;&lt;p&gt;An &lt;strong&gt;Online Judge (OJ)&lt;/strong&gt; is a web-based system used in competitive programming. It presents algorithmic problems, accepts code submissions in various languages, automatically compiles and runs the code against hidden test cases, and immediately tells the user whether their solution is correct, all without any human involvement in the grading process. Platforms like LeetCode and Codeforces are well-known examples.&lt;/p&gt;&#10;&lt;p&gt;Verwandlung Online Judge is a self-hostable, open-source OJ built for running your own contests or practice environment. Its main distinguishing feature at the time of release was cross-platform support: most open-source OJs were Linux-only due to their reliance on Linux-specific sandboxing APIs, whereas Verwandlung runs natively on both Windows and Linux.&lt;/p&gt;</description></item><item><title>TestZilla</title><link>https://www.infinitescript.com/project/testzilla/</link><pubDate>Thu, 01 Jan 2015 11:40:00 +0000</pubDate><guid>https://www.infinitescript.com/project/testzilla/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;&#10;&lt;p&gt;&lt;strong&gt;Crowd testing&lt;/strong&gt; is a software testing approach where a distributed group of real users, rather than an in-house QA team, tests a product across a wide variety of devices, operating systems, and usage scenarios. The goal is to surface bugs that structured internal testing tends to miss, at a fraction of the cost of a dedicated test lab.&lt;/p&gt;&#10;&lt;p&gt;TestZilla is an open-source crowd testing platform designed to connect two parties: &lt;strong&gt;hunters&lt;/strong&gt; (volunteer testers who find and report bugs) and &lt;strong&gt;developers&lt;/strong&gt; (product owners who register their software and receive the reports). Hunters earn points for accepted reports, creating a lightweight incentive loop that encourages quality submissions. The project started as a Spring MVC application and was later fully rebuilt using the Phalcon framework to support a faster continuous delivery workflow.&lt;/p&gt;</description></item><item><title>Medical Image Tagger</title><link>https://www.infinitescript.com/project/medical-image-tagger/</link><pubDate>Wed, 14 May 2014 11:21:55 +0000</pubDate><guid>https://www.infinitescript.com/project/medical-image-tagger/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;&#10;&lt;p&gt;&lt;strong&gt;Medical Image Tagger (MITagger)&lt;/strong&gt; is a collaborative web-based annotation platform developed with Harvard Medical School to accelerate semi-automatic tagging of biomedical figures. By combining NLP-extracted figure legends with the BioPortal Annotator REST API, the system recommends structured tags from major medical ontologies, reducing the manual effort of medical annotators by 20%.&lt;/p&gt;&#10;&lt;p&gt;The platform later evolved into a broader medical big data initiative, &lt;strong&gt;ShuYi Technology (数翼科技)&lt;/strong&gt;, which won the &lt;strong&gt;Silver Award&lt;/strong&gt; at the &lt;a href="https://rjxy.hfut.edu.cn/info/1024/2893.htm"&gt;&amp;ldquo;科蓝杯&amp;rdquo; 9th HFUT Student Entrepreneurship Competition&lt;/a&gt; in 2014.&lt;/p&gt;</description></item><item><title>Student Management System</title><link>https://www.infinitescript.com/project/student-management-system/</link><pubDate>Mon, 01 Apr 2013 11:21:00 +0000</pubDate><guid>https://www.infinitescript.com/project/student-management-system/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;&#10;&lt;p&gt;The Student Management System is a PHP/CodeIgniter web application built for the &lt;a href="http://rjxy.hfut.edu.cn"&gt;School of Software at Hefei University of Technology&lt;/a&gt; to replace a fully paper-based &lt;strong&gt;Comprehensive Evaluation&lt;/strong&gt; (综合测评): an annual multi-dimensional student assessment covering academic performance, peer review, extracurricular involvement, and community service, used to determine scholarship rankings.&#10;&lt;a href="https://rjxy.hfut.edu.cn/info/1016/1968.htm"&gt;According to school news coverage&lt;/a&gt;, what previously took over 5 days of manual tallying per cohort was reduced to microsecond-level computation.&#10;The project was selected as a &lt;strong&gt;&lt;a href="https://rjxy.hfut.edu.cn/info/1016/2274.htm"&gt;National Undergraduate Innovation Training Program&lt;/a&gt;&lt;/strong&gt; (国家级大学生创新训练计划项目) project in 2013.&lt;/p&gt;</description></item><item><title>HFUT Portals</title><link>https://www.infinitescript.com/project/hfut-portals/</link><pubDate>Fri, 01 Mar 2013 12:00:00 +0000</pubDate><guid>https://www.infinitescript.com/project/hfut-portals/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;&#10;&lt;p&gt;As an undergraduate at Hefei University of Technology (HFUT), I designed and built official websites for several university departments, listed below with the period each stayed in service.&#10;Each site runs on a custom WordPress theme, which let department staff publish notices, news, and documents on their own through the familiar dashboard.&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;&lt;a href="https://web.archive.org/web/20150924182207/http://rjxy.hfut.edu.cn/"&gt;School of Software&lt;/a&gt;&lt;/strong&gt; (软件学院): March 2013 to March 2018&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;&lt;a href="https://web.archive.org/web/20150922071240/http://dxsxlgh.hfut.edu.cn/welcome/"&gt;Psychological Care Center&lt;/a&gt;&lt;/strong&gt; (大学生心理关爱网): from November 2013&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;&lt;a href="https://web.archive.org/web/20170525165515/http://jwbjxb.hfut.edu.cn/"&gt;Office of Academic Affairs&lt;/a&gt;&lt;/strong&gt; (教学办公室): August 2014 to October 2020&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;&lt;a href="https://web.archive.org/web/20220501134934/http://rwyszjyzx.hfut.edu.cn/"&gt;Center for Humanities and Quality Education&lt;/a&gt;&lt;/strong&gt; (人文与素质教育中心): July 2015 to May 2022&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;&lt;a href="https://web.archive.org/web/20211122111355/http://jiwei.hfut.edu.cn/"&gt;Discipline Inspection and Supervision Office&lt;/a&gt;&lt;/strong&gt; (纪委办公室、监察处): August 2014 to November 2021&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The Psychological Care Center was &lt;a href="https://news.hfut.edu.cn/info/1011/15603.htm"&gt;launched&lt;/a&gt; after nearly a year of planning.&#10;It combines five sub-sites for students, parents, peer counselors, faculty, and mental health professionals, and was entered in the 6th National Top 100 University Websites (第六届全国高校百佳网站) contest run by the Ministry of Education&amp;rsquo;s China University Students Online.&lt;/p&gt;</description></item><item><title>Software QA System</title><link>https://www.infinitescript.com/project/software-qa-system/</link><pubDate>Fri, 11 Nov 2011 11:30:00 +0000</pubDate><guid>https://www.infinitescript.com/project/software-qa-system/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;&#10;&lt;p&gt;The Software QA System is a network-based online examination suite developed for the Software Basis Knowledge Contest (软件基础知识竞赛), part of the 3rd &amp;ldquo;Software Festival&amp;rdquo; (软件文化节) hosted by the &lt;a href="http://rjxy.hfut.edu.cn"&gt;School of Software at Hefei University of Technology&lt;/a&gt;.&#10;Initially developed in 2011 and later refactored in 2013, the system is built using VB.NET with a Microsoft Access (.mdb) backend, and uses UDP sockets for real-time client–server communication.&lt;/p&gt;&#10;&lt;p&gt;The system was officially deployed at the contest held on March 25, 2012, replacing traditional paper-based exams for over 70 participants from multiple departments.&#10;The use of this custom-built software was &lt;a href="https://news.hfut.edu.cn/info/1017/23644.htm"&gt;highlighted in the university news&lt;/a&gt; as a demonstration of the school&amp;rsquo;s professional expertise.&lt;/p&gt;</description></item><item><title>Carbon Footprint Calculator</title><link>https://www.infinitescript.com/project/carbon-footprint-calculator/</link><pubDate>Tue, 16 Mar 2010 21:29:00 +0000</pubDate><guid>https://www.infinitescript.com/project/carbon-footprint-calculator/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;&#10;&lt;p&gt;&lt;strong&gt;Carbon Footprint Calculator&lt;/strong&gt; (碳足迹计算器) is a desktop application I built in March 2010 using Visual Basic 6.0, an early project I built in middle school. The goal was to help people understand the environmental cost of everyday activities and encourage concrete steps to reduce CO₂ emissions.&lt;/p&gt;&#10;&lt;p&gt;In 2010, the application won the &lt;strong&gt;second prize&lt;/strong&gt; in the Visual Basic Programming Contest in Hangzhou (杭州市VB程序设计大赛).&lt;/p&gt;&#10;&lt;p&gt;You can download the application from &lt;a href="https://cloud.haozhexie.com/s/KwCW/qnysry50"&gt;this link&lt;/a&gt;, but please note that the application is in &lt;strong&gt;Simplified Chinese&lt;/strong&gt;.&lt;/p&gt;</description></item></channel></rss>