Production incidents should not begin with an engineer rebuilding context by hand.Specialized agents gather operational evidence and return a traceable root cause for incident engineers to review.Each production trace feeds a controlled self-improvement loop that catches regressions and proposes engineer-approved changes.The underlying data platform handles 10B events a day and peaks near 250K per second without sacrificing freshness or recovery.
What I build- Automated Multi-Agent Diagnosis
- First-pass incident triage with evidence-backed root causes
- Controlled Agent Improvement Loop
- Production traces become regression signals and engineer-approved improvements
- High-Volume Data Infrastructure
- 10B events a day, peaks near 250K per second, with freshness and recovery
Weixiang Wan, Senior Software Engineer
Weixiang Wan
The AI agent era is already unfolding. I choose to grow with it and help shape what comes next.
A calm field of independent particles establishes the site’s visual language.
Weixiang Wan
The same living particle field drifts aside, resolves into a personal portrait, and releases the incident signal that begins the first project.
Weixiang Wan
The living portrait holds while the personal introduction arrives one line at a time.
Automated Incident Diagnosis
Production incidents rarely fail for lack of data. The real problem is that the right evidence is scattered across logs, metrics, code, and operational knowledge.
An alert sits beside four distinct evidence systems that have not yet been assembled into one investigation.
Automated Incident Diagnosis
I designed an agent diagnosis harness to assemble that context automatically. Specialized investigators work in parallel, while every tool call, claim, and recommendation remains bounded and traceable.
Five specialized investigators gather evidence in parallel and pass bounded findings into a shared synthesis path.
Automated Incident Diagnosis
This changes the engineer's role from rebuilding an investigation to reviewing an evidence-backed root cause and deciding what to do next.
Evidence becomes a causal chain, a next action, and a review-ready diagnosis without collapsing the stages into one opaque box.
Platform Observability and Agent Insights
A single incident can look isolated. I built a fleet-wide insights layer to reveal the platform pattern behind it.
One diagnosis expands into a fleet-level operating view.
Platform Observability and Agent Insights
Across more than 430 repositories, raw telemetry creates too much noise. I designed an insights layer that scores suspicious signals and enriches them with code and quality evidence.
Anomalous repository clusters gather supporting evidence.
Platform Observability and Agent Insights
Not every anomaly deserves an investigation. The workflow I built promotes only evidence-backed patterns into diagnosis, turning fleet telemetry into focused engineering work.
The fleet view filters anomalies into the diagnosis harness.
Agent Observability and Evaluation Flywheel
Automation stops scaling when every answer still needs an engineer’s judgment. I saw evaluation—not generation—as the next bottleneck to solve.
A diagnosis unfolds into its complete production trace.
Agent Observability and Evaluation Flywheel
A production trace becomes useful when it can be judged against a stable reference. I built an evaluation flywheel that compares real agent behavior with a maintained golden dataset through independent evaluators.
Independent evaluators compare traces with a stable reference set.
Agent Observability and Evaluation Flywheel
Evaluation matters only when its failures lead somewhere useful. The flywheel separates regressions from latent failures, proposes targeted changes, and sends only exceptions back to engineering review.
Failures separate into targeted improvement paths.
Agent Observability and Evaluation Flywheel
Approved changes return to production automatically. Each new trace becomes evidence for the next evaluation cycle, so the system keeps improving without hiding the final decision from engineers.
Approved changes return as a higher-level evaluation cycle.
Declarative Data Orchestration
Production workflows become difficult to scale when every team describes them differently. I worked toward one declarative definition that captures structure, dependencies, and execution needs.
Distributed activity condenses into one portable definition.
Declarative Data Orchestration
A definition only matters if teams can trust how it runs. I helped build an engine that validates the same specification, generates serial or parallel Airflow DAGs, and carries the workflow across distributed compute backends.
A validated definition expands into a portable execution graph.
Declarative Data Orchestration
The abstraction became shared infrastructure for more than 350 live pipelines across 38 engineering teams. Teams could ship workflows without maintaining their own orchestration framework.
One engine becomes a fleet of live data workflows.
High Volume Streaming and Batch Data
Before data can support real-time or analytical workloads, it needs a reliable way into the platform. I helped shape a shared ingestion path that gives every event the same controlled entry point.
Independent events converge into a shared ingestion plane.
High Volume Streaming and Batch Data
The same path supports both continuous decisions and large analytical workloads. I contribute to streaming and batch pipelines that process about ten billion events a day and sustain peaks near 250 thousand events per second.
The same event population becomes streaming and batch paths at scale.
High Volume Streaming and Batch Data
Streaming and batch solve different timing problems, but downstream teams should not have to reconcile two versions of the truth. I help bring both paths into analytics-ready datasets while protecting throughput, freshness, and completeness.
Streaming and batch paths settle into one analytical layer.
Streaming Reliability
At this scale, reliability problems rarely begin everywhere at once. A single hot partition can spread into Kafka lag, scheduling delay, and stale data downstream.
One overloaded partition distorts the entire streaming path.
Streaming Reliability
Symptoms appear in different parts of the system, so the loudest signal is not always the root cause. I correlate partition activity, throughput, batch duration, and executor health to find where pressure actually begins.
Operational signals converge on the true source of pressure.
Streaming Reliability
Once the source is clear, recovery becomes a controlled balancing problem. I tune partition distribution, consumer parallelism, and intake rate until work moves evenly through the system again.
Work redistributes until every partition moves at a stable rate.
Automated Data Revision and Recovery
Streaming systems can remain operational while their data quietly becomes incomplete. The real failure appears when the destination no longer matches the source.
Source and destination datasets reveal a silent discrepancy.
Automated Data Revision and Recovery
Rebuilding everything would be expensive and would hide the actual failure boundary. I redesigned the revision workflow to detect the discrepancy, identify affected partitions, and produce a targeted recovery plan.
Only affected partitions are selected for revision.
Automated Data Revision and Recovery
The recovery path rebuilds only the missing data instead of replaying the entire system. That reduces manual recovery work while protecting downstream completeness.
Targeted data returns until both datasets match.
Supply Chain Intelligence
Logistics events are individually ordinary; their sequence is what reveals shipment risk. I designed a batch intelligence pipeline that prepares those events for an ML risk model and turns its predictions into actionable signals.
Logistics events gather into scheduled batch computation.
Supply Chain Intelligence
A prediction is only useful when it reaches the people who can act on it. The pipeline classifies shipment risk and delivers the result to downstream applications where teams decide what needs attention.
Processed shipments separate into actionable risk groups.
Multi Region Cloud Resilience
Regional resilience is difficult to trust when the secondary environment is assembled differently from the primary. I designed reusable infrastructure modules that reproduce the same platform foundations across both regions.
One operating region gains a structurally identical counterpart.
Multi Region Cloud Resilience
Matching infrastructure is only valuable when workloads can move safely during a regional failure. The design preserves networking, identity, compute, and observability as traffic transitions to the secondary region.
Work moves to the secondary region as the primary region dissolves.
FlowPulse
Agents can accelerate investigation without becoming the source of truth.
FlowPulse emerges from the shared project constellation.
Research Engine
Agent research should expose what was found, what failed, and what remains unanswered.
Research Engine emerges from the shared project constellation.
ClawBound
Model creativity can sit inside a runtime whose task, context, tools, and side effects remain explicit.
ClawBound emerges from the shared project constellation.
Let’s talk
I like systems that make difficult operations easier to understand, recover, and trust.
The shared particle population resolves into a personal signature.