I'm a Senior Product Manager with five years of experience building AI/ML platforms and infrastructure at scale, from early-stage startups to FAANG. I've shipped at every stage: scrappy zero-to-one builds where I was writing requirements, running user interviews, and partnering with the CTO in the same week, all the way to billion-transaction platforms spanning 150+ devices across 40+ facilities in six countries.
I'm the PM who figures out how to get it done, and who knows the difference between a problem that needs AI and a problem that needs fixing. I use AI to simplify and move fast, but the target is always the opportunity, never the technology. I stay on top of frontier research to know what just became possible, where it applies, and how to actually use it. And when I build, I design around the corner: for what models will make possible in six months, not just what they can do today, so the product gets better as the models do. On nights and weekends, I'm building multi-agent AI systems, orchestration layers, eval frameworks, and agentic tools, because the problems at the frontier are too interesting to just manage from a roadmap.
I was single-threaded owner of three charters inside Amazon Shipping: agentic AI platforms, shipment creation products, and hardware/edge compute. Together they spanned a network moving 1B+ parcels a year across six countries. I wrote the multiyear strategies that decided where capex, headcount, and engineering effort went, and owned 150+ devices across 40+ global facilities end to end: vendor relationships, sensor-to-cloud pipelines, the works. $440M+ in P&L impact across the three.
Nobody was taking the space, and the opportunity was obvious. 35+ account managers were losing hours a day stitching together Salesforce, SQL queries and dispute-tracking tools, or waiting on other teams for information they couldn't pull themselves. Every one of those hours was an hour not spent with customers. I wrote the org's multiyear agentic AI strategy and found others who were as interested in solving it as I was: an SDM and a group of engineers who wanted to work on AI. We built it on existing LLM infrastructure.
The architectural bet: build the foundation first, with agentic workflows and multi-agent orchestration underneath, so the same platform could be leveraged across domains instead of rebuilt for each one. That foundation is what moved teams from reactive to proactive: instead of responding to disputes and defects after they surfaced, the system surfaced them first.
Trust was the hard part, not the model. Eval discipline, accuracy thresholds and human-in-the-loop guardrails shipped before rollout: access boundaries, output transparency, escalation paths. An account manager won't act on a recommendation they can't audit. That's what cleared the enterprise trust bar.
Five organizations had worked on this for years without cracking it. The blocker was never difficulty. Everyone was solving against constraints that had been inherited rather than verified, so the solutions on the table were aimed at the wrong problem. I went back to first principles, found those limits weren't actually real, and proved it with data in a half-day test I ran myself. It had been solvable the entire time.
Then I expanded what the system was allowed to be. The proprietary billing ML already running on installed hardware could serve as a real-time decision layer instead of passive capture. That reframe turned a multi-year problem into the org's priority initiative, with VP sponsorship across all five organizations.
I used AI to do the integration homework before asking anyone for their time. Rather than booking meetings and asking other organizations to scope the work for me, I used AI to understand their tech stacks and dependencies well enough to draft a design that would actually fit them. By the first conversation I had a validated approach to react to instead of a blank page, which is why alignment took weeks rather than quarters. The $200M+ is the labor savings alone; it doesn't count the downstream improvements to claims, support, and everything else that gets better when packages stop kicking out.
The sensor network existed to bill customers. I turned it into a data layer for computer vision. 150+ dimensioning and scanning devices across 40+ facilities were built for billing capture. Nobody had decided where ML inference should actually run, or how sensor data should feed model inputs. I wrote the 3-year edge compute strategy that got VP approval: define on-device versus cloud inference by failure mode, and redirect infrastructure investment accordingly.
Owning it meant owning all of it: vendor relationships, reliability, and the pipeline from sensor capture to cloud storage. I authored the strategy for deploying Amazon-owned Outpost PCs as edge compute, shifting control away from third-party vendors and laying the groundwork for on-device ML at the edge.
Hardware drifts silently, which is the expensive part. I designed the automated calibration system (ConTest) so devices stopped quietly falling out of spec, recovering 1,100+ labor hours a year that had been going into manual testing.
Hardware had never had a single owner. Operations, Vendor Management and Site Management each touched the scanner network, so scan failures got absorbed as a cost of doing business rather than treated as a problem with a root cause. That fragmentation was the opportunity: the network needed one holistic strategy, so I expanded my charter to take hardware and build it.
The root cause of the root-cause mystery was the measurement layer. I went to CVG9, ONT5, and ACY9 to watch the failures happen. That's where I found it: the manufacturer's error-code reporting had been misconfigured for months, so every prior analysis had been built on wrong data. Calibration schedules didn't exist at all.
Then I A/B tested the fixes instead of rolling them out on faith, and adapted OEE from manufacturing into a metric the team had never owned, with baselines, thresholds and dashboards so the gains would outlive my involvement. Every mis-scan is a package that arrives late and a shipper who stops trusting the network, so the payoff was a better delivery experience and $60M in avoided cost, with zero added headcount and zero capex. Set a Black Friday network record along the way: 136K packages sorted in a single day.
My manager was aligned on compressing the launch to hit the quarter. I said no, and showed the math. Enterprise shippers needed programmatic shipment creation and label purchase across six countries, under pressure to cut corners on onboarding to hit a date. I pushed back with data from prior launches where thin documentation had turned into permanent support burden and slow adoption, and proposed a phased alternative: SA-supported shippers first, then scale the process.
The decision landed partway against me. I committed fully anyway, and ran the first integrations myself (BarkBox, American Eagle, Funko, Ipsy) to pressure-test the docs rather than hand them off. The standards I'd argued for became the onboarding template for every geography launch that followed.
Then the same template opened a new market. B2B expansion unblocked PO-number support and multi-parcel shipments, opening a $41B addressable segment. Label customization across the EU cut picker/packer time 60%.
Nobody assigned me to teach the org AI. I built the tools that made it easy to start anyway. The team wasn't behind on AI because people didn't want it. The on-ramp didn't exist. On my own time I built Cedric, a library of AI personas and templates so people could start from a working example instead of a blank prompt, and FlashBrief, an MCP-server tool that turns project noise into a digestible summary.
Then I coached without judgment, sharing my own failed prompts alongside the wins. Normalizing experimentation mattered more than any documentation. Both tools became shared org resources, and I ended up the most active AI practitioner on the team. Never the goal; just what happens when the on-ramp is good enough that people actually use it.
The diagnostic flow had 25% too many steps. I found that out by asking clinicians, not by guessing. Piction's core differentiator was its dermatology diagnosis flow, and it needed an overhaul. I ran 40+ clinician interviews and re-derived the workflow from scratch rather than patching the one that existed. That's how the redesign cut a quarter of the steps out.
Then we had to tell the story to investors. I owned iOS onboarding end to end: research, prototyping, engineering handoff. I also co-wrote the Series A product narrative with the CEO. The round closed. This is also where I got the conviction that AI products have to work reliably in high-stakes environments, because in this one a wrong answer is a misdiagnosis.
I went looking for the one metric that showed whether patients actually trusted the product. The messaging funnel moves through four stages: sent, read, opened, acted on. It predicted both platform LTV and hospital compliance, and it wasn't instrumented well enough to manage. I shipped SMS API integration and read-receipt instrumentation to make it measurable across 250K+ clinical users.
Then I audited feature parity across iOS, Android, and web, and rather than just filing the findings, turned two gaps into funded roadmap projects and shipped 80% of the rest myself.
Think of these as the components that make the system run. Most were stress-tested in production at Amazon scale; the rest come from systems I build and operate myself.
AI-native systems I build on nights and weekends. Some open source, some private.
The AI I actually run my life through, over iMessage, like texting a person. Six months of daily use, with 1,038 dated work-log entries, one for every meaningful change.
It runs on a library of skills and specialized subagents, each one narrow and good at a single thing, composed as the task demands, with an orchestration layer deciding how much compute each task actually deserves. At the frontier end it does research synthesis: tracking new papers and industry news, working out what's worth acting on and where the opportunities are. In the middle it does the product work: PRDs, company analysis, market research, interview coaching, and surfacing role opportunities worth a serious look. And at the small end it does things that are just useful. Send it a restaurant from Instagram and it pulls the details, synthesizes them, writes them into my Google Maps, and remembers what I liked, so the next recommendation lands closer.
Memory is where most of the work went. A knowledge graph of roughly 195,000 entities sits underneath, with temporal triples tracking not just what is true but when it was true, and contradiction detection that flags when two facts disagree. Around that: full-text search, a local embedding model for semantic recall, hybrid and multi-strategy retrieval, incremental indexing, weekly consolidation, and a decay pass so old context ages out instead of accumulating forever. Six scheduled jobs keep it current.
And it corrects itself. Every night it reviews the corrections I gave it that day and turns each one into a durable lesson injected into future runs. It doesn't just remember what I told it. It remembers to stop making the same mistake.
Home robots are shipping now. 1X's NEO, Figure 03. The hard part was never the folding. A housekeeper already knows how to clean a house. What they don't know on day one is your house: which desk is whose, where things go, what never gets touched. You walk them through it once, in about thirty minutes, and they're operational.
That's the problem I'm interested in. Fleet-scale humanoids will ship with a general world model and no local context, and every home is different. Capability generalizes. Context doesn't. So how does a system get personalized fast, from the smallest possible amount of input, without the owner manually demonstrating every preference? What is the least you have to say before it starts to feel like yours?
HomeGraph is where I'm working that out: a schema and a test harness for the foundational world-modeling data a system needs to localize itself to one specific home. I scoped it to text deliberately, holding perception constant so the knowledge-acquisition problem could be measured on its own. Whether a robot can see the cleats is a solved-enough problem. Whether it knows whose they are, and what to do when it doesn't, is not.
doctor command.