Portrait of Venkat Kumar Laxmi Kanth Nemala

Venkat Kumar Laxmi Kanth Nemala

NYU Computer Engineering graduate and Research Assistant at the AI4CE Lab, working on world models, multimodal learning, robotic planning, and 3D scene understanding.

I build machine-learning systems that connect advanced research with real-world robotics, autonomous systems, and document intelligence.

World Models & 3D Scene Understanding

Research at the intersection of world models, robotic planning, and visual perception.

DINO-WM

Implemented a VQ-VAE-based quantized variant of DINO-WM using pretrained DINOv2 visual representations for zero-shot planning and robotic manipulation, achieving a 4% performance improvement.

Visit DINO-WM ↗

Multiview Scene Graph (MSG)

Developed and improved the Multiview Scene Graph model for multiview perception and 3D scene understanding using the large-scale ScanNet++ image and video dataset, improving MSG edge IoU by 7%.

Created a large-scale training dataset containing more than 1,000 scenes from the ScanNet++ dataset for MSG model development and evaluation.

Visit MSG ↗

Education

New York University

New York, NY

Master of Science in Computer Engineering

CGPA: 3.8/4.0

August 2024 – May 2026

Vignan’s Institute of Information Technology

Visakhapatnam, India

Bachelor of Technology in Computer Science and Engineering

CGPA: 3.7/4.0

August 2019 – May 2023

Skills

Programming

Python, Go, JavaScript, TypeScript, C/C++, Java

Frontend

React, HTML, CSS, JavaScript/TypeScript

Backend / APIs

Node.js, FastAPI, Flask, REST APIs, Distributed Systems, Asynchronous Services

ML Frameworks

PyTorch, TensorFlow, JAX, torchvision, Core ML

GPU / Low-Level

CUDA, OpenCL, TensorRT, TVM, XLA

Testing / DevOps

Unit Testing, End-to-End Testing, CI/CD, Docker, Kubernetes, AWS Lambda, AWS Cognito, Amazon S3, AWS SageMaker, GCP Vertex AI, Linux

Databases

PostgreSQL, MongoDB, Redis, SQL

Tools / Workflow

Git, Terraform, Ansible, MLflow, Airflow, Prometheus, Grafana, Weights & Biases, Agile/Scrum Methodologies

Experience

Research Assistant NYU AI4CE Lab

September 2024 – Present · New York, NY

  • Developed and improved the Multiview Scene Graph (MSG) model using the large-scale ScanNet++ image and video dataset for multiview visual perception and 3D scene understanding, improving MSG edge IoU by 7%.
  • Created a large-scale training dataset containing more than 1,000 scenes from the ScanNet++ dataset for MSG model development and evaluation.
  • Implemented a VQ-VAE-based quantized variant of DINO-WM using pretrained DINOv2 visual representations for zero-shot planning and robotic manipulation, achieving a 4% performance improvement.
  • Explored diffusion-based generative models using Diffusion Transformers (DiT) for structured scene representation and visual synthesis in 3D environments.

Teaching Assistant | DS-GA 1008 Deep Learning New York University · Prof. Yann LeCun

January 2025 – May 2025 · New York, NY

  • Mentored students in deep learning architectures and multimodal machine-learning systems for Prof. Yann LeCun’s graduate Deep Learning course.
  • Supported course instruction, technical discussions, assignment evaluation, and student guidance throughout the semester.

Section Leader | DS-GA 1011 Natural Language Processing New York University

September 2024 – December 2024 · New York, NY

  • Led NLP sections and mentored students through practical labs covering high-performance computing (HPC), reinforcement learning from human feedback (RLHF), and information retrieval.
  • Helped more than 200 students across the teaching roles understand deep learning, NLP, and multimodal ML concepts and implementations.

Member of Technical Staff – I Alphastream.ai

August 2023 – August 2024 · Bengaluru, India

  • Engineered an end-to-end AutoML pipeline for dataset preparation, model training, evaluation, and deployment, reducing retraining time from 9 days to 3 days and improving extraction accuracy by 9%; led dataset curation and Label Studio annotation workflows with a team of 7 engineers.
  • Fine-tuned and integrated the Co-DETR model to extract complex hierarchical patterns from tables, addressing 95% of prior model limitations and analyzing model robustness across evaluation datasets.
  • Applied multimodal LayoutLMv3 for document visual-content understanding, identifying table text and images in PDFs and achieving a 14% accuracy improvement.

Python Developer Intern Alphastream.ai

December 2022 – July 2023 · Bengaluru, India

  • Developed an NLP-based parser integrated with Amazon Textract OCR to mask sensitive information, achieving 99% data-masking accuracy and improving extraction performance by 13%.
  • Built a RAG pipeline for news-article processing, benchmarking LLMs including LLaMA, Qwen, and Mistral and achieving 97% summarization accuracy.
  • Designed a personalized article-recommendation and ad-targeting pipeline based on user reading behavior, improving overall user engagement by 30%.

Projects

Predictive Steering with I-JEPA

Pretrained I-JEPA on Waymo and CARLA datasets and fine-tuned with ADL-JEPA for label-efficient steering-angle prediction, achieving 99.3% accuracy and validating in CARLA.

Python · PyTorch · CARLA · Waymo · ADL-JEPA

View on GitHub ↗

Multimodal Action Localization

Enhanced VSLNet with LaViLa vision-language embeddings, stacked encoders, and vision-enhancer attention for Ego4D action localization, achieving 12.14 mIoU.

PyTorch · Ego4D · Omnivore · LaViLa

View on GitHub ↗

Research Paper Summarizer

Built model-serving and monitoring components for a scientific-paper summarization system using transformer models, arXiv data, Flask APIs, and production observability tooling.

BART · Flask · Prometheus · Grafana · MLOps

View on GitHub ↗

Data Structures & Algorithms

Implemented foundational data structures and algorithms in C, including queues, stacks, linked lists, heaps, tries, trees, graph traversal, Dijkstra, Prim, and Huffman coding.

C · Data Structures · Graph Algorithms

View on GitHub ↗

Camp Grounds

Developed a campground web application with user authentication, reviews, and RESTful CRUD operations for creating and managing campground listings.

JavaScript · MongoDB · Express.js · Node.js

View on GitHub ↗

Human Posture Detection

Enhanced real-time posture detection by fine-tuning MoveNet with TensorFlow, achieving 93.73% accuracy, and built a user-friendly CNN-powered interface for improved accessibility.

Python · Flask · TensorFlow.js · MoveNet · CNN

View on GitHub ↗

Traffic Signs Recognition

Created a convolutional neural-network application that classifies traffic signs, with model experimentation in Jupyter and a Flask-based prediction interface.

CNN · TensorFlow · Keras · Flask · Jupyter

View on GitHub ↗

Highlights

Research Impact

Improved MSG edge IoU by 7% and DINO-WM planning and manipulation performance by 4% through research at NYU AI4CE Lab.

Teaching at NYU

Served separately as Teaching Assistant for Prof. Yann LeCun’s Deep Learning course and NLP Section Leader, leading labs on HPC, RLHF, and information retrieval while mentoring more than 200 students.

Contact

Email: vn2263@nyu.edu

Phone: +1 (347) 798-7171

Location: New York City

GitHub: github.com/nvklaxmikanth

LinkedIn: linkedin.com/in/nvklaxmikanth