Ruby (Yuru) XuSoftware Engineer
I build the systems that products and models run on.

Backend platforms, cloud infrastructure and LLM serving — five domains since 2016, from distributed microservices built from scratch to foundation-model MLOps at Google X. Twice promoted to lead the team building them.
Munich, GermanyOpen to Software Engineer RolesRequires visa sponsorshipRemote or relocation
Five domains, one job
- 2016 — 2018Security platformsDistributed microservices, task queues, automated scanning
- 2018 — 2019Game backendsAWS and Tencent Cloud, CI/CD, load balancing, DB clustering
- 2019 — 2022Real-time 3DUnity, depth cameras, Lidar, commercial installations
- 2022 — 2024Computer visionFoundation models, few-shot detection, Vertex AI
- 2025 — nowLLM infrastructureServing, routing, observability, performance tuning

01GMI Cloud2025 — 2026
PerfLab
A monitoring and auto-tuning service for the model serving fleet, built to replace manual health checks and hand-tuned deployment configuration.
2h → <30minto deploy, tune and test a model
I led the project. A monitoring service, delivered as a working demo in two weeks, probes model health continuously and surfaces a failing endpoint inside a minute. Wiring it into our inference engine turned configuration tuning into a systematic sweep across serving frameworks and their environment variables, instead of a person trying values by hand.
Project lead
How it was builtClose
- Problem
- Endpoint failures surfaced whenever somebody happened to notice them, and arriving at a workable deployment configuration for a new model took over two hours of manual testing.
- Ownership
- I led the project, from the first demo through its integration with the inference engine.
- Approach
- A monitoring service in FastAPI and Node.js, delivered as a working demo in two weeks, continuously probing online model health and performance. I then integrated it with our internal inference engine to automate configuration tuning — systematically testing service frameworks such as vLLM and SGLang along with their environment variables to determine the optimal parameters rather than guessing at them.
- Impact
- Endpoint failures are now detected within one minute. Deployment, tuning and performance testing fell from more than two hours per model to under thirty.
02GMI Cloud2025 — 2026
LLM Proxy & API Gateway
A traffic layer in front of multiple model backends, built so one unhealthy backend does not take the whole service down with it.
I led traffic governance and reliability for the gateway every request passes through. Routing follows live health scores and automated probes, failover and retry handling cover the backends that stop answering, and OpenTelemetry makes what each one is doing visible rather than inferred. I then moved the routing and scoring intelligence out of the Python proxy into a Go control plane.
Traffic governance & reliability lead
How it was builtClose
- Problem
- An LLM gateway sits on the critical path for every request. Its routing, its failure handling and its visibility into what the backends are actually doing determine whether a partial outage is absorbed or passed straight through to users.
- Ownership
- I led traffic governance and reliability engineering for the multi-vendor proxy and API gateway, and re-architected its routing and scoring intelligence into a Go control plane.
- Approach
- Dynamic routing with health scoring and automated probing, so traffic follows the backends that are actually healthy; failover, timeout and retry handling for the ones that are not; and observability across all of them. Built on Nacos, Redis and OpenTelemetry. The routing and scoring intelligence then moved from the Python proxy into a Go control plane: an endpoint scoring engine with a config-driven tier classifier and traffic weighting, and a latency-based capacity detector that catches endpoint overload dynamically.
- Impact
- High availability and day-to-day operability across multiple model backends, with the behaviour of each one visible rather than inferred, and overloaded endpoints caught dynamically rather than after the fact.
03Google X · Mineral.ai2022 — 2024
Project Dolphin
A foundation-model-native MLOps system, built to remove annotation as the bottleneck in computer-vision work.
10×image annotation throughput
My modeling work on an earlier project led to being made tech lead, and I took five engineers — three frontend, two backend — from an empty repository to a launched MVP in six months. A foundation model serves few-shot detection from custom containers on Vertex AI, proposing the labels a human would otherwise draw one at a time.
Tech lead & project lead, team of 5
How it was builtClose
- Problem
- Image annotation gated everything downstream. Models could not improve faster than humans could label the data they were trained on.
- Ownership
- I was promoted to tech lead and project leader on the strength of my modeling impact on an earlier project, then led a team of five — three frontend and two backend engineers — building the application from the ground up.
- Approach
- Smart labeling driven by foundation models: OWL-ViT served for few-shot and one-shot object detection from custom Docker containers, productionized behind APIs on Vertex AI. Pairing the foundation model with downstream tasks accelerated response times while improving accuracy. Architecture decisions were made by forming hypotheses and running controlled experiments, evaluating alternatives quantitatively rather than by argument.
- Impact
- The MVP launched within six months and improved image annotation efficiency tenfold.
04Shanghai IntuQ2016 — 2018
Vulnerability Management Platform
A security vulnerability management SaaS, built from an empty repository so scanning, tracking and reporting run as one distributed system.
6 monthsfrom first engineering job to leading the team
My first engineering job, and where the systems instinct came from. I built the platform with a three-person team — asset discovery, a Celery task-queue tier with workers, and reporting on top of it — and was promoted to team lead within six months to head the work.
Promoted to team lead, team of 3
How it was builtClose
- Problem
- Vulnerability management only helps if the scanning, the asset inventory and the reporting are one system. Run as separate manual steps, findings age faster than anyone can act on them.
- Ownership
- I was promoted to team lead within six months of joining, and led a three-person team building the platform.
- Approach
- A microservice architecture built from the ground up in Python and Django, with Celery task queues and workers driving automated scanning, scripts modelled on white-hat techniques gathering asset data from RPC sources, and pandas-backed reporting and charts on top. I handled server operations and infrastructure across the distributed deployment.
- Impact
- A working SaaS platform, and the distributed-systems foundation everything I have built since rests on.

How I work
Leadership has arrived twice by promotion. I was made team lead within six months at my first engineering job, and at Google X the impact of my modeling work led to being made tech lead of Project Dolphin — five engineers, from an empty repository to a launched MVP in six months. I design cloud-native architecture and decide it empirically, running controlled experiments rather than arguing from preference, and I sit with stakeholders as readily as with the team.
LLM inference & serving
Keeping model endpoints fast and available when individual backends degrade — dynamic routing, health scoring, failover, and tuning the frameworks that serve them.
- vLLM
- SGLang
- FastAPI
- Nacos
- Redis
- OpenTelemetry
ML systems & MLOps
Getting models out of research and into something people can depend on — foundation models containerized, served behind APIs, and built into platforms around them.
- Vertex AI
- OWL-ViT
- Docker
- BigQuery
- Beam / Flume
- Kubernetes
Distributed backend & cloud
The layer underneath both, absorbing load and failure — microservices built from the ground up, task queues and workers, architecture decided by experiment.
- Python
- Celery
- RabbitMQ
- Django
- Google Cloud
- AWS
Track record
- 2026.1 — PresentX-LABProject Manager & Requirements Lead, Software Engineer
- 2025.5 — 2026.9GMI CloudSoftware Engineer, Large Language Model
- 2024.5 — 2024.9Cerbo AISoftware Engineer, Large Language Model
- 2022.6 — 2024.5Google X · Mineral.aiMachine Learning Software Engineer → Tech Lead
- 2020.9 — 2022.6Shanghai Ang Di AdvertisingUnity 3D Engineer
- 2019.8 — 2020.7Freelance · AALabUnity 3D Engineer, Blender Developer
- 2018.10 — 2019.7Shanghai Yi Zuan Network TechnologySoftware Engineer
- 2016.6 — 2018.9Shanghai IntuQ Information Security TechnologySoftware Engineer → Team Lead
Before machine learning: six years of real-time 3D and interactive installations wired to depth cameras and Lidar, backend infrastructure for high-traffic mobile games, and a distributed security platform I was promoted to lead.
Education
Master of Engineering, specialized in Machine Learning
Arizona State University2023.9 — 2025.7GPA 4.0 / 4.0
M.Sc. Data Science
XU Exponential University of Applied Sciences2025.10 — 2027.9In progress
B.Sc. Computer Science and Technology
Chengdu University of Information Technology2013.9 — 2017.6
Publication
Sparsity Limit to Prune Large Language Models for On-Device AI Assistants: Llama-2 as an Example
Liu, B.; Xu, Y.Preprints 2024
Full toolkitClose
- Languages
- Python (proficient) · Golang · C# · SQL, NoSQL, Bigtable
- Machine learning
- End-to-end model delivery, deployment and integration in cloud environments · n8n
- Backend
- Redis · RabbitMQ · Celery · Kubernetes · distributed systems
- Cloud
- Google Cloud (proficient) · AWS · Azure · Alibaba Cloud
- Platforms
- Linux (Ubuntu, CentOS) · macOS · Windows
- Methodologies
- Agile · CI/CD · microservices · containerization
Contact
If you are building something that has to hold up, I would like to hear about it.
Munich, GermanyOpen to Software Engineer RolesRequires visa sponsorshipRemote or relocation
