Anthony Ledesma, Staff AI Engineer

The run, in text

  1. 2016GoDaddy, Tier 3 Support, then Software Engineer I Nov 2016 to Oct 2019

    Started in enterprise hosting support. Taught myself to code.

    • Shipped React and Node features for apps serving 20M+ customers.
  2. 2019GoDaddy, Software Engineer Oct 2019 to Apr 2024

    The call: Smart Templates inference is too expensive. Where do you cut?

    • Switch to a cheaper model. A quick win on the bill, paid for in output quality.
    • Refactor the retrieval and generation pipeline (my call). More work up front, and the product keeps its quality bar.

    THE COST BEAST: It eats a little more of the budget with every request. How I beat it:

    1. Refactor retrieval. Retrieval, rebuilt.
    2. Refactor generation. Generation, rebuilt to match.
    3. Ship it and measure. Inference cost down 93%. Latency down 71%. The project that proved I could own an AI product end to end.
  3. 2024GoDaddy, Senior Software Engineer Apr 2024 to Apr 2026

    Built MCP-style agent tool layers before the protocol was standardized.

    The call: Agents need WordPress and platform operations. Vendor SDKs, or hand-roll the connectors?

    • Use the vendor SDKs. Faster start, and every SDK still has to clear security review.
    • Hand-roll the connectors (my call). No third-party SDK waiting on security review. The team owns every connector it ships.

    CONTEXT ROT: Six agent layers share one context, and any tool can fire on any turn. How I beat it:

    1. Split the prompt into interpretation layers. One domain per layer, with controlled handoffs between stages.
    2. Gate tools on conversation state. Each API and extraction service runs when the state calls for it, not on every turn.
    3. Route dispatch through structured output schemas. Schemas decide dispatch and what context moves between stages. Six agent layers, shipped on deadline for a high-visibility executive demo.
    • Worked across 5 engineering groups to codify and share AI patterns.
    • Cut CI/CD runtimes 83% (60+ min to 10 min).
  4. 2025McLean Forrester LLC (Contract), Senior AI Engineer Sep 2025 to Mar 2026

    The call: The output changes from run to run. How do releases get gated?

    • Pass or fail on a single run. Simple, and it flips on noise alone.
    • A tolerance band over N runs (my call). Golden baselines from real usage, each eval run N times to a distribution. More runs per release, and a gate that does not flip on noise.

    THE 39-POINT REGRESSION: An upstream integration changes what the model sees. Did quality hold? How I beat it:

    1. Measure the noise floor first. Every fixture run to a distribution before anything changes. The steady ones hold under two points of spread.
    2. Assert in three layers. Schema, property invariants, behavior. Then Welch's t-test and Glass's delta on the numbers, gating releases in CI.
    3. Compare the new build against baseline. A 39-point drop against a baseline spread of about 2 is roughly twenty sigma. It never shipped.
    • Took over an inherited Terraform migration on GCP and delivered full platform parity, with zero secrets in state.
  5. 2026Tracine, Open-source tooling for AI coding agents 2026

    The call: An eval arm comes back with zero spread. Refuse to compute, or compute and warn?

    • Refuse. Safe, and blind. A flat arm is a ceiling or a broken harness, and refusing tells you neither.
    • Compute it and name the constant arm (my call). eval-kit (alpha) computes it and names the flat arm. The verdict stands, with the warning first.

    THE NOISE: Four runs per arm. Shift +2.70, 95% CI [-0.84, +6.24], target 3. How I beat it:

    1. Read the interval, not the mean. +2.70 looks like a win. The interval holds zero, and it holds the 3-point target. p = 0.111.
    2. Size the run before trusting it. At the spread seen so far, settling a 3-point call takes 16 runs per arm. We have 4.
    3. Run 16 fresh per arm. At 16 the interval can no longer hold both zero and a 3-point shift. Whatever comes back is a call.
    • guard: allow, deny or ask on agent tool calls, before they run.
    • convo: every session and tool call, searchable.
  6. NOWTaskRay, Staff AI EngineerJun 2026 to now

    Applied LLM and agentic systems.

    WritingGithubLinkedin

Writing

All posts

Notes from building AI systems, from what an eval can actually tell you to the tooling I run agents with.

First posts are in review. Back soon.

Open source

tracine.dev