About

Tingde Liu

I am a robotics and AI engineer with an M.Sc. in Mechatronics and Robotics from Leibniz Universität Hannover (LUH). My current research centers on vision-language navigation (VLN) and 3D Large Language Models (3DLLM) — and the broader question of what it would actually take for a robot to understand and act in the world the way we do. I am interested in genuine embodied intelligence, not just systems that appear to navigate, but ones that truly reason about space, language, and intention.

Why I Created This Repository

This repository is my working space for collecting, organizing, and sharing the ideas that shape my research. It brings together survey posts, paper notes, and long-form reflections on embodied AI, with a focus on VLN, 3D Large Language Models (3DLLM), robotics, and the systems that connect perception, reasoning, and action.

I also want it to be a friendly open-source space where people feel welcome to read, contribute, discuss, and help improve the ideas here together.

I created it to make my learning process visible and reusable: a place to track what I read, what I build, and how my understanding changes over time. Instead of keeping those notes scattered across documents and bookmarks, I wanted a public archive that can grow with my work and stay useful to anyone exploring the same space.

The goal is simple: turn ongoing research into something structured, searchable, and worth revisiting.

Research Interests

These experiences converged around a set of questions I keep returning to:

  • How can language models reason meaningfully about 3D space?
  • What does it take for a robot to navigate using natural instructions?
  • How do we bridge the gap between simulation and real-world perception?
  • How do we build a Robot OS — a coherent harness that integrates perception, memory, planning, and action into a unified system?
  • How can agentic AI give robots something closer to genuine agency: not just executing instructions, but forming intentions, adapting plans, and acting with purpose?

In practice this means I work on Vision-Language Navigation (VLN), 3DLLM, agentic robot systems, and the infrastructure that makes embodied intelligence real.

Technical Skills

Languages: Python, C++, MATLAB Frameworks & Systems: PyTorch, ROS2, LangChain, CUDA, Vision-Language & Multimodal: CLIP, BLIP, LLaVA, Vision-Language Modeling, Multimodal LLMs 3D Vision & Perception: LiDAR processing, PCL, 3D Gaussian Splatting, Yolo Robotics: Robot Perception, Motion Planning, INS, SLAM, Imitation Learning, Reinforcement Learning, Policy Learning Simulation & Tools: Claude Code, Docker, Git, Gazebo, Isaac Sim, Habitat, Openclaw


Continuously learning and exploring the infinite possibilities of AI and Robotics!