Skip to content

Ruby (Yuru) XuSoftware Engineer

I build the systems that products and models run on.

Ruby Xu, photographed in Munich, Germany
Munich, Germany

Backend platforms, cloud infrastructure and LLM serving — five domains since 2016, from distributed microservices built from scratch to foundation-model MLOps at Google X. Twice promoted to lead the team building them.

Munich, GermanyOpen to Software Engineer RolesRequires visa sponsorshipRemote or relocation

Five domains, one job

  • 2016 — 2018Security platformsDistributed microservices, task queues, automated scanning
  • 2018 — 2019Game backendsAWS and Tencent Cloud, CI/CD, load balancing, DB clustering
  • 2019 — 2022Real-time 3DUnity, depth cameras, Lidar, commercial installations
  • 2022 — 2024Computer visionFoundation models, few-shot detection, Vertex AI
  • 2025 — nowLLM infrastructureServing, routing, observability, performance tuning
Deep blue open sea seen past a branch of yellow-green euphorbia.
PROBEENDPOINTSCONFIG SWEEP<1 MINvLLM / SGLang · ENV VARS↳ OPTIMALDEPLOY → TUNE → TEST2H+ → <30 MIN

01GMI Cloud2025 — 2026

PerfLab

A monitoring and auto-tuning service for the model serving fleet, built to replace manual health checks and hand-tuned deployment configuration.

2h → <30minto deploy, tune and test a model

I led the project. A monitoring service, delivered as a working demo in two weeks, probes model health continuously and surfaces a failing endpoint inside a minute. Wiring it into our inference engine turned configuration tuning into a systematic sweep across serving frameworks and their environment variables, instead of a person trying values by hand.

Project lead

How it was built
Problem
Endpoint failures surfaced whenever somebody happened to notice them, and arriving at a workable deployment configuration for a new model took over two hours of manual testing.
Ownership
I led the project, from the first demo through its integration with the inference engine.
Approach
A monitoring service in FastAPI and Node.js, delivered as a working demo in two weeks, continuously probing online model health and performance. I then integrated it with our internal inference engine to automate configuration tuning — systematically testing service frameworks such as vLLM and SGLang along with their environment variables to determine the optimal parameters rather than guessing at them.
Impact
Endpoint failures are now detected within one minute. Deployment, tuning and performance testing fell from more than two hours per model to under thirty.
REQUESTSGATEWAYMODEL BACKENDSHEALTH SCORINGDYNAMIC ROUTINGTIMEOUT / RETRYFAILOVERBACKEND A0.98BACKEND BFAILBACKEND C0.94OPENTELEMETRY · NACOS · REDISONE BACKEND DOWN ≠ OUTAGE

02GMI Cloud2025 — 2026

LLM Proxy & API Gateway

A traffic layer in front of multiple model backends, built so one unhealthy backend does not take the whole service down with it.

I led traffic governance and reliability for the gateway every request passes through. Routing follows live health scores and automated probes, failover and retry handling cover the backends that stop answering, and OpenTelemetry makes what each one is doing visible rather than inferred. I then moved the routing and scoring intelligence out of the Python proxy into a Go control plane.

Traffic governance & reliability lead

How it was built
Problem
An LLM gateway sits on the critical path for every request. Its routing, its failure handling and its visibility into what the backends are actually doing determine whether a partial outage is absorbed or passed straight through to users.
Ownership
I led traffic governance and reliability engineering for the multi-vendor proxy and API gateway, and re-architected its routing and scoring intelligence into a Go control plane.
Approach
Dynamic routing with health scoring and automated probing, so traffic follows the backends that are actually healthy; failover, timeout and retry handling for the ones that are not; and observability across all of them. Built on Nacos, Redis and OpenTelemetry. The routing and scoring intelligence then moved from the Python proxy into a Go control plane: an endpoint scoring engine with a config-driven tier classifier and traffic weighting, and a latency-based capacity detector that catches endpoint overload dynamically.
Impact
High availability and day-to-day operability across multiple model backends, with the behaviour of each one visible rather than inferred, and overloaded endpoints caught dynamically rather than after the fact.
UNLABELLEDFOUNDATION MODELANNOTATEDOWL-ViTFEW-SHOT / ONE-SHOTSERVED ON VERTEX AIBY HANDSMART LABELLING10× THROUGHPUT

03Google X · Mineral.ai2022 — 2024

Project Dolphin

A foundation-model-native MLOps system, built to remove annotation as the bottleneck in computer-vision work.

10×image annotation throughput

My modeling work on an earlier project led to being made tech lead, and I took five engineers — three frontend, two backend — from an empty repository to a launched MVP in six months. A foundation model serves few-shot detection from custom containers on Vertex AI, proposing the labels a human would otherwise draw one at a time.

Tech lead & project lead, team of 5

How it was built
Problem
Image annotation gated everything downstream. Models could not improve faster than humans could label the data they were trained on.
Ownership
I was promoted to tech lead and project leader on the strength of my modeling impact on an earlier project, then led a team of five — three frontend and two backend engineers — building the application from the ground up.
Approach
Smart labeling driven by foundation models: OWL-ViT served for few-shot and one-shot object detection from custom Docker containers, productionized behind APIs on Vertex AI. Pairing the foundation model with downstream tasks accelerated response times while improving accuracy. Architecture decisions were made by forming hypotheses and running controlled experiments, evaluating alternatives quantitatively rather than by argument.
Impact
The MVP launched within six months and improved image annotation efficiency tenfold.
ASSET DISCOVERYTASK QUEUEREPORTSRPC SOURCERPC SOURCERPC SOURCEWORKERSPYTHON · DJANGO · CELERY · REDIS · RABBITMQ · MONGODBMICROSERVICES, BUILT FROM AN EMPTY REPOSITORY

04Shanghai IntuQ2016 — 2018

Vulnerability Management Platform

A security vulnerability management SaaS, built from an empty repository so scanning, tracking and reporting run as one distributed system.

6 monthsfrom first engineering job to leading the team

My first engineering job, and where the systems instinct came from. I built the platform with a three-person team — asset discovery, a Celery task-queue tier with workers, and reporting on top of it — and was promoted to team lead within six months to head the work.

Promoted to team lead, team of 3

How it was built
Problem
Vulnerability management only helps if the scanning, the asset inventory and the reporting are one system. Run as separate manual steps, findings age faster than anyone can act on them.
Ownership
I was promoted to team lead within six months of joining, and led a three-person team building the platform.
Approach
A microservice architecture built from the ground up in Python and Django, with Celery task queues and workers driving automated scanning, scripts modelled on white-hat techniques gathering asset data from RPC sources, and pandas-backed reporting and charts on top. I handled server operations and infrastructure across the distributed deployment.
Impact
A working SaaS platform, and the distributed-systems foundation everything I have built since rests on.
A wooded headland dropping into a hazy sea, two small islands offshore.

How I work

Leadership has arrived twice by promotion. I was made team lead within six months at my first engineering job, and at Google X the impact of my modeling work led to being made tech lead of Project Dolphin — five engineers, from an empty repository to a launched MVP in six months. I design cloud-native architecture and decide it empirically, running controlled experiments rather than arguing from preference, and I sit with stakeholders as readily as with the team.

LLM inference & serving

Keeping model endpoints fast and available when individual backends degrade — dynamic routing, health scoring, failover, and tuning the frameworks that serve them.

  • vLLM
  • SGLang
  • FastAPI
  • Nacos
  • Redis
  • OpenTelemetry

ML systems & MLOps

Getting models out of research and into something people can depend on — foundation models containerized, served behind APIs, and built into platforms around them.

  • Vertex AI
  • OWL-ViT
  • Docker
  • BigQuery
  • Beam / Flume
  • Kubernetes

Distributed backend & cloud

The layer underneath both, absorbing load and failure — microservices built from the ground up, task queues and workers, architecture decided by experiment.

  • Python
  • Celery
  • RabbitMQ
  • Django
  • Google Cloud
  • AWS

Track record

  • 2026.1 — PresentX-LABProject Manager & Requirements Lead, Software Engineer
  • 2025.5 — 2026.9GMI CloudSoftware Engineer, Large Language Model
  • 2024.5 — 2024.9Cerbo AISoftware Engineer, Large Language Model
  • 2022.6 — 2024.5Google X · Mineral.aiMachine Learning Software Engineer → Tech Lead
  • 2020.9 — 2022.6Shanghai Ang Di AdvertisingUnity 3D Engineer
  • 2019.8 — 2020.7Freelance · AALabUnity 3D Engineer, Blender Developer
  • 2018.10 — 2019.7Shanghai Yi Zuan Network TechnologySoftware Engineer
  • 2016.6 — 2018.9Shanghai IntuQ Information Security TechnologySoftware Engineer → Team Lead

Before machine learning: six years of real-time 3D and interactive installations wired to depth cameras and Lidar, backend infrastructure for high-traffic mobile games, and a distributed security platform I was promoted to lead.

Education

  • Master of Engineering, specialized in Machine Learning

    Arizona State University2023.9 — 2025.7GPA 4.0 / 4.0

  • M.Sc. Data Science

    XU Exponential University of Applied Sciences2025.10 — 2027.9In progress

  • B.Sc. Computer Science and Technology

    Chengdu University of Information Technology2013.9 — 2017.6

Full toolkit
Languages
Python (proficient) · Golang · C# · SQL, NoSQL, Bigtable
Machine learning
End-to-end model delivery, deployment and integration in cloud environments · n8n
Backend
Redis · RabbitMQ · Celery · Kubernetes · distributed systems
Cloud
Google Cloud (proficient) · AWS · Azure · Alibaba Cloud
Platforms
Linux (Ubuntu, CentOS) · macOS · Windows
Methodologies
Agile · CI/CD · microservices · containerization

Contact

If you are building something that has to hold up, I would like to hear about it.

Munich, GermanyOpen to Software Engineer RolesRequires visa sponsorshipRemote or relocation