<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://zhengthomastang.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://zhengthomastang.github.io/" rel="alternate" type="text/html" /><updated>2026-06-11T22:19:46-07:00</updated><id>https://zhengthomastang.github.io/feed.xml</id><title type="html">Dr. Zheng (Thomas) Tang</title><subtitle>Senior Deep Learning Engineer at NVIDIA</subtitle><author><name>Dr. Zheng (Thomas) Tang</name><email>tangzhengthomas@gmail.com</email></author><entry><title type="html">Cosmos 3, VANTAGE-Bench, and TAR at Computex 2026</title><link href="https://zhengthomastang.github.io/posts/2026/06/blog-post-1/" rel="alternate" type="text/html" title="Cosmos 3, VANTAGE-Bench, and TAR at Computex 2026" /><published>2026-06-01T00:00:00-07:00</published><updated>2026-06-01T00:00:00-07:00</updated><id>https://zhengthomastang.github.io/posts/2026/06/blog-post-1</id><content type="html" xml:base="https://zhengthomastang.github.io/posts/2026/06/blog-post-1/"><![CDATA[<p>At <strong>Computex 2026</strong>, NVIDIA announced <strong><a href="https://blogs.nvidia.com/blog/cosmos-3-physical-ai-open-world-foundation-model/">Cosmos 3</a></strong>, an open world foundation model for physical AI that brings vision reasoning, multimodal generation, and action prediction together for robots, autonomous vehicles, smart spaces, and vision AI agents.</p>

<p>One of the most meaningful parts for me was seeing two benchmarks I supported, <strong><a href="https://vantage-bench.org/">VANTAGE-Bench</a></strong> and <strong><a href="https://www.aicitychallenge.org/2026-track3/">Traffic Anomaly Reasoning (TAR)</a></strong>, highlighted as part of the Cosmos 3 release story. Both benchmarks focus on the kind of operational video understanding that matters for physical AI: fixed cameras, long videos, small but important events, spatial-temporal reasoning, and explanations that go beyond object detection.</p>

<p><strong>VANTAGE-Bench</strong> evaluates vision-language models on real-world fixed-camera footage across warehouse/logistics, transportation, and smart-space domains. The benchmark contains 3,346 image and video assets with 35,027 expert-curated annotations, organized around semantic, spatial, temporal, and spatio-temporal understanding. I helped prepare VANTAGE-Bench as a <strong>NeurIPS 2026 competition effort</strong>, supporting the benchmark and leaderboard framing around the “Infrastructure AI Gap”: how well VLMs can produce physically grounded insights from fixed infrastructure cameras under realistic deployment constraints.</p>

<p><strong>TAR</strong> is the foundation of <strong>AI City Challenge 2026 Track 3: Anomalous Events in Transportation</strong>, part of the 10th AI City Challenge at ECCV 2026. TAR moves traffic anomaly evaluation from binary detection toward multi-task reasoning, with 44,040 training annotations across 10 task types covering 3,670 CCTV transportation videos. The track asks models to detect, reason about, and explain traffic anomalies through question answering, temporal reasoning, causal linkage, scene description, and video summarization. I supported the launch alignment for TAR as the Track 3 benchmark, helping keep the evaluation and leaderboard direction focused and ready for the Cosmos 3 public launch.</p>

<p>Together, these efforts show a broader shift in video AI evaluation: from recognizing visible objects to understanding what happened, why it happened, and what might happen next. They also connect to my broader MetroAI direction on multimodal AI for cities, transportation, and infrastructure, where video summarization, event understanding, and embedding-based video search are becoming central tools for operational intelligence. It was exciting to see that work connected to the Cosmos 3 launch and to the larger push toward physical AI systems that can reason over real-world environments.</p>

<p align="center">
  <img src="https://zhengthomastang.github.io/images/Cosmos3_Computex_Keynote.jpg" alt="Jensen Huang's Computex keynote slide announcing NVIDIA Cosmos 3" style="width: 850px;" />
</p>]]></content><author><name>Dr. Zheng (Thomas) Tang</name><email>tangzhengthomas@gmail.com</email></author><category term="work" /><category term="research" /><summary type="html"><![CDATA[NVIDIA's Cosmos 3 release at Computex 2026 highlighted physical AI benchmarks including VANTAGE-Bench and Traffic Anomaly Reasoning (TAR). I supported VANTAGE-Bench preparation as a NeurIPS 2026 competition effort and TAR launch alignment for AI City Challenge 2026 Track 3.]]></summary></entry><entry><title type="html">10th AI City Challenge Accepted as an ECCV 2026 Workshop</title><link href="https://zhengthomastang.github.io/posts/2026/04/blog-post-1/" rel="alternate" type="text/html" title="10th AI City Challenge Accepted as an ECCV 2026 Workshop" /><published>2026-04-12T00:00:00-07:00</published><updated>2026-04-12T00:00:00-07:00</updated><id>https://zhengthomastang.github.io/posts/2026/04/blog-post-1</id><content type="html" xml:base="https://zhengthomastang.github.io/posts/2026/04/blog-post-1/"><![CDATA[<p>I am excited to share that the <strong>10th AI City Challenge</strong> has been accepted as a half-day workshop at <strong><a href="https://eccv.ecva.net/Conferences/2026/Workshops">ECCV 2026</a></strong> in Malmo, Sweden. The official ECCV workshop list includes our workshop under the title <strong>“Towards Sim2Real Transfer and Unified Reasoning: 10th AI City Challenge”</strong>.</p>

<p>This milestone continues a decade of AI City Challenge workshops and competitions, bringing together researchers and practitioners working on computer vision, multimodal AI, intelligent transportation, smart cities, and large-scale video analytics.</p>

<p>The 2026 challenge focuses on <strong>Sim2Real transfer</strong>, <strong>unified reasoning</strong>, and <strong>smart-city video AI</strong> across six tracks:</p>

<ol>
  <li><strong>Multi-Camera 3D Perception</strong>, for Sim2Real multi-object tracking across camera networks.</li>
  <li><strong>Transportation Safety Understanding and Captioning</strong>, for captioning and VQA around pedestrian-centric risk.</li>
  <li><strong>Anomalous Events in Transportation</strong>, for detecting, localizing, and explaining traffic anomalies.</li>
  <li><strong>Text-Based Person Re-Identification</strong>, for natural-language retrieval by appearance and actions.</li>
  <li><strong>Generative Traffic Video Forecasting</strong>, for future-frame generation from video history and text descriptions.</li>
  <li><strong>Cross-City Object Detection</strong>, a Milestone Systems Hafnia track for geographic generalization.</li>
</ol>

<p>I am grateful to continue helping organize this community as it moves from detection-centric perception toward richer reasoning, prediction, and deployment-aware evaluation for real-world infrastructure.</p>

<p align="center">
  <img src="https://zhengthomastang.github.io/images/AICity2026_ECCV_Workshop.png" alt="10th AI City Challenge ECCV 2026 workshop poster with six challenge tracks" style="width: 760px;" />
</p>]]></content><author><name>Dr. Zheng (Thomas) Tang</name><email>tangzhengthomas@gmail.com</email></author><category term="work" /><category term="research" /><summary type="html"><![CDATA[The 10th AI City Challenge was accepted as a half-day workshop at ECCV 2026 in Malmo, Sweden, continuing a decade of benchmarking real-world computer vision and AI for transportation, smart cities, and large-scale video analytics.]]></summary></entry><entry><title type="html">Video Analytics AI Agents at GTC 2026</title><link href="https://zhengthomastang.github.io/posts/2026/03/blog-post-1/" rel="alternate" type="text/html" title="Video Analytics AI Agents at GTC 2026" /><published>2026-03-19T00:00:00-07:00</published><updated>2026-03-19T00:00:00-07:00</updated><id>https://zhengthomastang.github.io/posts/2026/03/blog-post-1</id><content type="html" xml:base="https://zhengthomastang.github.io/posts/2026/03/blog-post-1/"><![CDATA[<p>At <strong>GTC 2026</strong>, I was excited to support the NVIDIA booth demo <strong>“Turn Video Into Insights With Video Analytics AI Agents”</strong>, highlighting the <strong>NVIDIA AI Blueprint for Video Search and Summarization (VSS)</strong>.</p>

<p>The demo showed how VSS can help developers build visual AI agents that review events, understand physical context, and streamline decisions across large volumes of live and recorded video. The workflow brought together search, summarization, Q&amp;A, active alerts, VLM/CV/model implementation, and event review, with examples spanning smart cities, warehouses, and factories.</p>

<p>My main contribution focused on the <strong>video embedding search</strong> stack behind the demo. I helped prepare search models and datasets, finalized embedding models and person attribute search checkpoints, generated example queries and captions for warehouse and other video scenarios, and supported evaluation data preparation for demo clips and KPI validation. I also helped debug and validate key pieces of the real-time search path, including text embedding behavior, model packaging, and search-profile integration.</p>

<p>This work connected directly to the broader VSS search roadmap. In the <strong><a href="https://docs.nvidia.com/vss/latest/release-notes.html#vss-3-1-0">VSS 3.1.0 release</a></strong>, the Search Agent Workflow added attribute search, multi-embedding fusion search, and a critic agent for reviewing search results. The release also updated the RT-CV microservice to support embedding generation for detected objects, including RADIO-CLIP and SigLIP2 embedding models.</p>

<p>For me, the most rewarding part was seeing several threads come together in one visible experience: video embeddings for event and activity search, object/person attribute search through vision-language embeddings, model packaging for deployment, and practical demo workflows that make large video archives feel searchable and useful.</p>

<p align="center">
  <img src="https://zhengthomastang.github.io/images/GTC_2026_VSS_Booth_Demo.jpg" alt="GTC 2026 VSS booth demo: Turn Video Into Insights With Video Analytics AI Agents" style="width: 750px;" />
</p>]]></content><author><name>Dr. Zheng (Thomas) Tang</name><email>tangzhengthomas@gmail.com</email></author><category term="work" /><category term="tech" /><summary type="html"><![CDATA[At GTC 2026, our VSS booth demo showed how video analytics AI agents can turn large volumes of live and recorded video into searchable, actionable insights. I contributed to the video embedding search components behind the demo and to the search features later highlighted in the VSS 3.1.0 release.]]></summary></entry><entry><title type="html">UrbanAI at NeurIPS 2025: Advancing Multi-Camera Tracking &amp;amp; Multimodal Spatial AI with NVIDIA Metropolis</title><link href="https://zhengthomastang.github.io/posts/2025/12/blog-post-1/" rel="alternate" type="text/html" title="UrbanAI at NeurIPS 2025: Advancing Multi-Camera Tracking &amp;amp; Multimodal Spatial AI with NVIDIA Metropolis" /><published>2025-12-07T00:00:00-08:00</published><updated>2025-12-07T00:00:00-08:00</updated><id>https://zhengthomastang.github.io/posts/2025/12/blog-post-1</id><content type="html" xml:base="https://zhengthomastang.github.io/posts/2025/12/blog-post-1/"><![CDATA[<p>I was honored to present at the <a href="https://urbanai2025.github.io/"><strong>UrbanAI Workshop</strong></a> at <a href="https://neurips.cc/"><strong>NeurIPS 2025</strong></a>, where we shared NVIDIA’s recent advancements in <strong>multi-camera 3D perception</strong>, <strong>cloud-native tracking workflows</strong>, and <strong>multimodal spatial AI</strong>.</p>

<p>We highlighted the nine-year evolution of the <strong>AI City Challenge</strong>, which now spans multi-camera 3D perception, traffic safety reasoning, warehouse spatial intelligence, and fisheye object detection—reflecting the growing needs of smart cities and intelligent infrastructure.</p>

<p>Our talk also introduced NVIDIA’s <strong>cloud-native, streaming multi-camera tracking workflow</strong>, combining DeepStream-based perception, MTMC fusion, real-time location tracking (RTLS), and a full <strong>Sim2Deploy</strong> pipeline powered by Omniverse, TAO Toolkit, and Metropolis microservices.</p>

<p>We shared progress on <strong>MCBLT</strong>, a BEV-based multi-camera 3D detection and hierarchical GNN tracking framework that achieves state-of-the-art results on AI City Challenge and WildTrack benchmarks.</p>

<p>Finally, we presented <strong>Sparse4D</strong>, an end-to-end multi-camera 3D perception model that jointly predicts 3D bounding boxes, tracking, velocity, and ReID embeddings—ranking #1 on the AICity’25 vision-only leaderboard.</p>

<p>It was great connecting with researchers from Google Mobility AI, DeepMind, Columbia University, and many others working to build the next generation of intelligent urban systems. Looking forward to continued collaboration across the UrbanAI and AI City Challenge communities.</p>

<p align="center">
  <img src="https://zhengthomastang.github.io/images/UrbanAI_NeurIPS_2025.png" alt="Urban AI Workshop at NeurIPS 2025" style="width: 750px;" />
</p>]]></content><author><name>Dr. Zheng (Thomas) Tang</name><email>tangzhengthomas@gmail.com</email></author><category term="work" /><category term="research" /><summary type="html"><![CDATA[At the NeurIPS 2025 UrbanAI Workshop, we presented nine years of progress from the AI City Challenge and introduced NVIDIA’s latest advancements in multi-camera 3D perception, cloud-native tracking workflows, and end-to-end spatial AI models like MCBLT and Sparse4D—supporting next-generation smart city applications.]]></summary></entry><entry><title type="html">Vision AI Agents for Foxconn Digital Twins Featured at NVIDIA GTC DC 2025 &amp;amp; Hosting the 9th AI City Challenge at ICCV 2025</title><link href="https://zhengthomastang.github.io/posts/2025/10/blog-post-1/" rel="alternate" type="text/html" title="Vision AI Agents for Foxconn Digital Twins Featured at NVIDIA GTC DC 2025 &amp;amp; Hosting the 9th AI City Challenge at ICCV 2025" /><published>2025-10-28T00:00:00-07:00</published><updated>2025-10-28T00:00:00-07:00</updated><id>https://zhengthomastang.github.io/posts/2025/10/blog-post-1</id><content type="html" xml:base="https://zhengthomastang.github.io/posts/2025/10/blog-post-1/"><![CDATA[<p>I’m honored that our work was featured in <a href="https://www.nvidia.com/gtc/dc/keynote/"><strong>Jensen Huang’s keynote at NVIDIA GTC DC 2025</strong></a>, where we showcased a new generation of <strong>vision AI agents</strong> operating inside a <strong>Foxconn factory digital twin</strong>.<br />
Built on <strong>NVIDIA Metropolis</strong>, <strong>Cosmos</strong>, and <strong>Omniverse</strong>, these agents monitor large industrial spaces from an overhead perspective, continuously analyzing activity, detecting anomalies, and assisting Foxconn engineers with real-time safety and operational insights.</p>

<p>This demo illustrates the future of <strong>Physical AI</strong>: unifying simulation, multimodal perception, and large-scale intelligence to power safer, more efficient, and more autonomous industrial environments. Grateful to our teammates across <strong>Metropolis</strong>, <strong>Robotics</strong>, and <strong>Omniverse</strong> who made this possible.</p>

<p>In parallel, I was also honored to represent <strong>NVIDIA Metropolis</strong> as host of the <a href="https://www.aicitychallenge.org/"><strong>9th AI City Challenge Workshop</strong></a> at <a href="https://iccv.thecvf.com/"><strong>ICCV 2025</strong></a> in Honolulu. This year, the challenge brought together <strong>245 teams from 15 countries</strong>, pushing forward research in:</p>

<ul>
  <li>Multi-camera 3D perception</li>
  <li>Traffic safety reasoning</li>
  <li>Warehouse spatial intelligence</li>
  <li>Edge-optimized fisheye detection</li>
</ul>

<p>The challenge datasets—including <a href="https://huggingface.co/datasets/nvidia/PhysicalAI-SmartSpaces">PhysicalAI-SmartSpaces</a> and <a href="https://woven-visionai.github.io/wts-dataset-homepage/">WTS Traffic Safety</a>—were downloaded <strong>nearly 400,000 times on Hugging Face</strong>, reflecting the global momentum behind large-scale spatial AI benchmarks.</p>

<p>Additionally, I was invited as a keynote speaker at the <a href="https://motchallenge.net/workshops/bmtt2025/"><strong>8th Workshop on Benchmarking Multi-Target Tracking (BMTT)</strong></a>, where I presented NVIDIA’s recent advances in:</p>

<ul>
  <li><a href="https://arxiv.org/abs/2412.00692"><strong>MCBLT (BEV-SUSHI)</strong></a>: hierarchical GNN-based multi-camera 3D tracking</li>
  <li><a href="https://catalog.ngc.nvidia.com/orgs/nvidia/teams/tao/models/sparse4d_rn101"><strong>Sparse4D</strong></a>: end-to-end multi-camera 3D perception</li>
  <li>Integration with <a href="https://github.com/dvl-tum/SUSHI"><strong>SUSHI GNN</strong></a> for scalable, long-range association</li>
</ul>

<p>Special thanks to collaborators <strong>Yizhou Wang</strong>, <strong>Sameer Pusegaonkar</strong>, and our partners in <strong>Laura Leal-Taixé’s group</strong>, whose contributions continue to push the boundaries of multi-camera 3D understanding.</p>

<p>Congratulations to all participants, winners, reviewers, and organizers who made this year’s workshops a success. Looking forward to driving the next wave of <strong>Spatial AI</strong>, <strong>digital twins</strong>, and <strong>large-space perception</strong> together with the community.</p>

<p align="center">
  <img src="https://zhengthomastang.github.io/images/AICity25_photo.jpg" alt="9th AI City Challenge Workshop at ICCV 2025" style="width: 750px;" />
</p>]]></content><author><name>Dr. Zheng (Thomas) Tang</name><email>tangzhengthomas@gmail.com</email></author><category term="work" /><category term="research" /><summary type="html"><![CDATA[Our vision AI agents for Foxconn’s factory digital twin were showcased in Jensen Huang’s GTC DC 2025 keynote, alongside hosting the 9th AI City Challenge Workshop at ICCV 2025 with 245 global teams advancing multi-camera 3D perception and spatial AI.]]></summary></entry><entry><title type="html">Releasing the Physical AI Smart Spaces Dataset for AI City Challenge 2024 &amp;amp; 2025</title><link href="https://zhengthomastang.github.io/posts/2025/03/blog-post-1/" rel="alternate" type="text/html" title="Releasing the Physical AI Smart Spaces Dataset for AI City Challenge 2024 &amp;amp; 2025" /><published>2025-03-18T00:00:00-07:00</published><updated>2025-03-18T00:00:00-07:00</updated><id>https://zhengthomastang.github.io/posts/2025/03/blog-post-1</id><content type="html" xml:base="https://zhengthomastang.github.io/posts/2025/03/blog-post-1/"><![CDATA[<p>I’m excited to share that our team at NVIDIA has officially released the <a href="https://huggingface.co/datasets/nvidia/PhysicalAI-SmartSpaces"><strong>Physical AI Smart Spaces</strong> dataset</a> as part of the <a href="https://blogs.nvidia.com/blog/open-physical-ai-dataset/"><strong>Open Physical AI Dataset</strong></a> initiative, announced during <strong>GTC 2025</strong>. This dataset supports <strong>Tracks 1</strong> of the <strong>AI City Challenge 2024 &amp; 2025</strong>, helping researchers push the boundaries of real-time, multi-camera spatial AI.</p>

<p>As a lead contributor, I helped drive the creation and integration of this large-scale synthetic dataset—designed specifically for <strong>multi-target multi-camera (MTMC) tracking</strong>, <strong>3D occupancy prediction</strong>, and <strong>smart infrastructure simulation</strong>. It is one of the most comprehensive synthetic datasets ever released for spatial AI, capturing warehouse, factory, and public space scenarios with fine-grained annotations and synchronized multi-view video.</p>

<p>Built with NVIDIA Omniverse and aligned with our Metropolis AI workflows, the dataset enables the development of robust computer vision systems for real-world deployment. From training deep learning models to validating digital twin applications, Physical AI Smart Spaces serves as a benchmark-ready resource for cutting-edge research and development.</p>

<p>We’re proud to empower the global research community with open tools and high-quality data. Whether you’re developing advanced trackers, occupancy estimators, or large-scale analytics systems, this dataset is built to accelerate your innovation.</p>

<p>Explore the dataset here:<br />
👉 <a href="https://huggingface.co/datasets/nvidia/PhysicalAI-SmartSpaces">huggingface.co/datasets/nvidia/PhysicalAI-SmartSpaces</a></p>

<p align="center">
  <img src="https://zhengthomastang.github.io/images/Open_Physical_AI_Dataset.jpg" alt="NVIDIA’s Open Physical AI Dataset" style="width: 750px;" /> 
</p>]]></content><author><name>Dr. Zheng (Thomas) Tang</name><email>tangzhengthomas@gmail.com</email></author><category term="work" /><category term="dataset" /><summary type="html"><![CDATA[Excited to announce the release of the Physical AI Smart Spaces dataset—developed as part of NVIDIA’s Open Physical AI Dataset initiative and featured at NVIDIA GTC 2025. This large-scale multi-camera 3D perception dataset supports AI City Challenge Tracks for 2024 and 2025, advancing research in multi-target multi-camera tracking, 4D occupancy, and digital twin applications.]]></summary></entry><entry><title type="html">Metropolis Spatial AI Powers Mega Omniverse Blueprint Featured at CES 2025</title><link href="https://zhengthomastang.github.io/posts/2025/01/blog-post-2/" rel="alternate" type="text/html" title="Metropolis Spatial AI Powers Mega Omniverse Blueprint Featured at CES 2025" /><published>2025-01-06T00:00:00-08:00</published><updated>2025-01-06T00:00:00-08:00</updated><id>https://zhengthomastang.github.io/posts/2025/01/blog-post-1</id><content type="html" xml:base="https://zhengthomastang.github.io/posts/2025/01/blog-post-2/"><![CDATA[<p>I’m incredibly proud to announce that our work on <strong>Metropolis Spatial AI</strong> was featured in the <strong>CES 2025 Keynote</strong> by NVIDIA CEO Jensen Huang, as a core component of the newly unveiled <a href="https://blogs.nvidia.com/blog/mega-omniverse-blueprint/"><strong>Mega Omniverse Blueprint</strong></a>. As the Tech Lead of this initiative within the Metropolis team, it’s a huge honor to see our technology playing a foundational role in enabling intelligent infrastructure at industrial scale.</p>

<p>The <strong>Mega Omniverse Blueprint</strong> is NVIDIA’s comprehensive framework for simulating, deploying, and scaling robot fleets in highly realistic, physics-informed digital twins. These virtual environments, powered by the NVIDIA Omniverse platform, allow organizations to design, test, and optimize robotics and AI systems before deploying them in the real world.</p>

<p>Our contribution through <strong>Metropolis Spatial AI agents</strong> includes real-time location systems (RTLS) powered by advanced multi-camera tracking, 3D perception using our BEV-SUSHI framework, and GNN-based tracking modules. These components enable accurate and scalable monitoring of autonomous agents—such as robots, vehicles, and humans—within massive industrial environments like warehouses and factories.</p>

<p>In collaboration with teams across NVIDIA, we integrated Metropolis Spatial AI into the Omniverse ecosystem alongside Isaac Sim™ and other robotics tools. The result is a tightly integrated digital twin solution for spatial intelligence—demonstrated in factory and warehouse scenarios during the CES keynote.</p>

<p>This marks a major step forward for real-time vision AI in digital twins, unlocking new levels of automation, safety, and operational efficiency. I’m beyond excited to see where this journey leads as we continue to push the boundaries of AI-powered infrastructure.</p>

<p align="center">
  <img src="https://zhengthomastang.github.io/images/Mega_Omniverse_Blueprint.jpg" alt="Metropolis Spatial AI Demo at CES'25" style="width: 750px;" /> 
</p>]]></content><author><name>Dr. Zheng (Thomas) Tang</name><email>tangzhengthomas@gmail.com</email></author><category term="work" /><category term="tech" /><summary type="html"><![CDATA[Excited to share that our work on Metropolis Spatial AI was showcased in the CES'25 Keynote by NVIDIA CEO Jensen Huang as part of the Mega Omniverse Blueprint—an end-to-end framework for simulating and deploying intelligent robot fleets. This milestone highlights the critical role of multi-camera tracking and real-time location systems in the future of industrial automation, smart factories, and digital twins.]]></summary></entry><entry><title type="html">Winning Best Presentation and Best Poster Awards at NTECH 2024</title><link href="https://zhengthomastang.github.io/posts/2025/01/blog-post-1/" rel="alternate" type="text/html" title="Winning Best Presentation and Best Poster Awards at NTECH 2024" /><published>2025-01-02T00:00:00-08:00</published><updated>2025-01-02T00:00:00-08:00</updated><id>https://zhengthomastang.github.io/posts/2025/01/blog-post-1</id><content type="html" xml:base="https://zhengthomastang.github.io/posts/2025/01/blog-post-1/"><![CDATA[<p>I am excited to share that our team received both the <strong>Best Presentation Award</strong> and the <strong>Best Poster Award</strong> in the <strong>Imaging and Computer Vision Track</strong> at <strong>NTECH 2024</strong>.</p>

<p>The <strong>Best Presentation Award</strong> recognized <strong>“BEV-SUSHI: Multi-Target Multi-Camera 3D Object Detection and Tracking in Bird’s-Eye View”</strong>, a collaboration with NVResearch-DVL. The work was authored by Yizhou Wang, Tim Meinhardt, Orcun Cetintas, Chris Yang, Sameer Satish Pusegaonkar, Benjamin Missaoui, Sujit Biswas, Zheng Tang, and Laura Leal-Taixe.</p>

<p>The <strong>Best Poster Award</strong> recognized <strong>“From MTMC to RTLS: Towards Real-Time, Scalable Multi-Camera Object Tracking and Localization”</strong>, authored by Zheng Tang, Ganapathy Seshadri Cadungude Aiyer, Akshay Agrawal, Shuo Wang, Sameer Satish Pusegaonkar, Kartikay Thakkar, and Sujit Biswas.</p>

<p>Both projects reflect our broader effort to move multi-target multi-camera tracking from offline perception pipelines toward real-time location systems for large spaces, industrial automation, and spatial AI.</p>

<p align="center">
  <img src="https://zhengthomastang.github.io/images/NTECH_2024_Best_Poster_Award.jpg" alt="NTECH 2024 Best Poster Award" style="width: 520px;" />
</p>]]></content><author><name>Dr. Zheng (Thomas) Tang</name><email>tangzhengthomas@gmail.com</email></author><category term="work" /><category term="award" /><summary type="html"><![CDATA[Our NTECH 2024 submissions received both the Best Presentation Award and Best Poster Award in the Imaging and Computer Vision Track, recognizing our work on BEV-SUSHI and real-time, scalable multi-camera object tracking and localization.]]></summary></entry><entry><title type="html">Announcing the Official Release of the Metropolis Multi-Camera Tracking AI Workflow</title><link href="https://zhengthomastang.github.io/posts/2024/07/blog-post-1/" rel="alternate" type="text/html" title="Announcing the Official Release of the Metropolis Multi-Camera Tracking AI Workflow" /><published>2024-07-02T00:00:00-07:00</published><updated>2024-07-02T00:00:00-07:00</updated><id>https://zhengthomastang.github.io/posts/2024/07/blog-post-1</id><content type="html" xml:base="https://zhengthomastang.github.io/posts/2024/07/blog-post-1/"><![CDATA[<p>I am thrilled to share that our pioneering work, the Metropolis Multi-Camera Tracking AI Workflow, was prominently featured in NVIDIA GTC’24 Keynote by Jensen Huang. As the Tech Lead of this groundbreaking project, it is an honor to announce that our innovative solution is now officially available.</p>

<p>Our AI-powered multi-camera tracking system is designed to accelerate the development of vision AI applications, providing a comprehensive solution for measuring and managing infrastructure and operations across large spaces.</p>

<p>Imagine a world where factories operate automatically with maximum safety and efficiency, retail spaces are optimized for a superior shopper experience, and public areas like hospitals, airports, and highways are safer and more streamlined. Multi-camera tracking allows for accurate object tracking and activity measurement across multiple cameras and spaces, enabling effective monitoring and management.</p>

<p>NVIDIA’s customizable multi-camera tracking workflow offers a validated path to production, eliminating months of development time. The solution includes state-of-the-art AI models pretrained on real and synthetic datasets, customizable for various use cases. The workflow spans from simulation to analytics and integrates NVIDIA’s cutting-edge tools, including Isaac SIM™, Omniverse™, TAO, and DeepStream. It features real-time video streaming modules and is built on a scalable, cloud-native microservices architecture. With expert support and the latest product updates through NVIDIA AI Enterprise, this workflow accelerates vision AI projects without additional costs, apart from infrastructure and tool licenses.</p>

<p>Applications of multi-camera tracking include enhancing manufacturing and warehouse operations by optimizing routes for autonomous robots, equipment, and workers. It also improves retail store layouts by analyzing customer navigation to maximize sales and revenue, and enhances in-hospital patient care by providing continuous monitoring for safety and security.</p>

<p>Our workflow includes the entire development pipeline—from data generation to model training to application development—helping developers build complex vision AI applications for large spaces. Key components include creating 3D digital twins of real-world environments using NVIDIA Omniverse, simplifying agent simulation with NVIDIA Isaac Sim, streamlining model training with NVIDIA TAO Toolkit, and using NVIDIA Metropolis microservices for modular, cloud-native application building blocks.</p>

<p>Jumpstart the development of multi-camera vision AI applications with NVIDIA’s end-to-end workflow, from Omniverse simulation for synthetic data generation to TAO for streamlined model development to Metropolis microservices for modular, cloud-native application building blocks.</p>

<p>Join us in this exciting journey towards transforming how we monitor and manage large spaces with cutting-edge AI technology.</p>

<p align="center">
  <img src="https://zhengthomastang.github.io/images/AI_Powered_Multi_Camera_Tracking.png?raw=true" alt="Photo" style="width: 750px;" /> 
</p>]]></content><author><name>Dr. Zheng (Thomas) Tang</name><email>tangzhengthomas@gmail.com</email></author><category term="work" /><category term="tech" /><summary type="html"><![CDATA[Announcing the official release of the Metropolis Multi-Camera Tracking AI Workflow, featured in NVIDIA GTC'24 Keynote by Jensen Huang. This innovative solution accelerates the development of vision AI applications for large spaces, enhancing safety, efficiency, and management across various industries. Leveraging NVIDIA's cutting-edge tools, this workflow offers a validated path to production, customizable AI models, and comprehensive support, enabling seamless development from simulation to deployment. Join us in transforming infrastructure and operations with advanced AI technology.]]></summary></entry><entry><title type="html">Advancing AI with NVIDIA Omniverse at CVPR 2024: AI City Challenge Highlights</title><link href="https://zhengthomastang.github.io/posts/2024/06/blog-post-2/" rel="alternate" type="text/html" title="Advancing AI with NVIDIA Omniverse at CVPR 2024: AI City Challenge Highlights" /><published>2024-06-17T00:00:00-07:00</published><updated>2024-06-17T00:00:00-07:00</updated><id>https://zhengthomastang.github.io/posts/2024/06/blog-post-1</id><content type="html" xml:base="https://zhengthomastang.github.io/posts/2024/06/blog-post-2/"><![CDATA[<p>As the lead organizer of the AI City Challenge at the Computer Vision and Pattern Recognition (CVPR) 2024 conference, I’m excited to share our progress, particularly with the creation of the largest indoor synthetic dataset using NVIDIA Omniverse.</p>

<p>The AI City Challenge draws over 700 teams from nearly 50 countries to develop AI models for improving operational efficiency in various physical settings. This year, NVIDIA Omniverse played a pivotal role, providing datasets for tasks like retail, warehouse management, and intelligent traffic systems.</p>

<p>In large indoor spaces like factories and warehouses, AI models require extensive data for training, which is often time-consuming and costly to collect manually. To overcome this, we used physically based simulations and digital twins created with NVIDIA Omniverse. These virtual environments generate synthetic data essential for training AI models to operate effectively and safely.</p>

<p>For this year’s AI City Challenge, we focused on the Multi-Camera Person Tracking track, the most popular with over 400 teams. NVIDIA provided a dataset with 212 hours of 1080p video at 30 frames per second, covering 90 scenes across six virtual environments, including warehouses, retail stores, and hospitals. These scenes, created in Omniverse, simulated nearly 1,000 cameras and featured around 2,500 digital human characters, allowing teams to test and refine their AI models accurately.</p>

<p>Our challenge saw global collaboration with ten prestigious institutions, including the Australian National University, Johns Hopkins University, and the Emirates Center for Mobility Research. These partnerships highlight the worldwide effort to advance AI technologies for smart cities and industrial automation.</p>

<p>Generative physical AI, combining reinforcement learning in simulated environments with high-fidelity physics-based simulation, is set to transform infrastructure automation and robotics. NVIDIA’s latest innovations, like the Omniverse Cloud Sensor RTX microservices, will speed up the development of fully autonomous systems by providing accurate sensor simulation in realistic virtual environments.</p>

<p>At CVPR 2024, NVIDIA Research will present over 50 papers on generative physical AI with applications in autonomous vehicle development and robotics. Highlights include unified 6D pose estimation and tracking, digital twin creation of unknown articulated objects, and customizable dataset generation via simulation.</p>

<p>For those interested in these advancements, NVIDIA offers a free standard license for Omniverse and extensive resources to get started. Join the Omniverse community on platforms like Instagram, Medium, LinkedIn, and Discord to stay updated and connect with other developers and researchers.</p>

<p>Together, we are driving AI innovation, creating smarter, safer, and more efficient environments for all.</p>

<p align="center">
  <img src="https://zhengthomastang.github.io/images/NVIDIA_Advances_Physical_AI_at_CVPR_With_Largest_Indoor_Synthetic_Dataset.jpg?raw=true" alt="Photo" style="width: 750px;" /> 
</p>]]></content><author><name>Dr. Zheng (Thomas) Tang</name><email>tangzhengthomas@gmail.com</email></author><category term="work" /><category term="tech" /><summary type="html"><![CDATA[As the lead organizer of the AI City Challenge at CVPR, I'm excited to highlight our progress with NVIDIA Omniverse, which provided the largest indoor synthetic dataset for over 700 teams from nearly 50 countries. This dataset, essential for developing AI models to improve efficiency in retail, warehouse management, and traffic systems, included 212 hours of video across 90 virtual environments. Our global collaboration with ten prestigious institutions underscores the effort to advance AI for smart cities and automation. NVIDIA's innovations, like Omniverse Cloud Sensor RTX, will further accelerate autonomous system development. Join the Omniverse community to stay updated and connected.]]></summary></entry></feed>