Emilio Cignetti

Senior Full-Stack Engineer

Italian and Argentine citizen

2025 to 2026Professional work

A production LLM pipeline that finds what slows engineers down

Development transcripts, scrubbed of secrets and personal data, read by Claude into typed findings. The confident ones arrive as pre-filled tickets.

A tool that reads engineers' transcripts has two ways to fail. It can send something it should never have sent, or it can produce findings nobody acts on. The second one kills it quietly.

  • Anthropic API
  • Claude
  • structured output
  • Go
  • Node.js
  • OAuth 2.0 device flow
  • Slack
  • Jira

All work

Engineering organizations are full of friction nobody reports. A build step everyone runs twice because it fails the first time. A script that does not exist, so eleven people wrote eleven versions of it. A failure mode that costs twenty minutes and is never worth the meeting it would take to raise.

None of that reaches a retrospective, because each instance is too small to mention and nobody holds the whole picture. The transcripts of engineers working, however, hold all of it.

So I built a pipeline that reads them: session hooks captured development transcripts, a worker scrubbed secrets and personal data, Claude returned typed findings with confidence scores naming missing tooling and repeated failure patterns, and findings above a threshold arrived in Slack as pre-filled Jira tickets for a human to accept or reject.

It shipped inside my team’s developer tooling, written in Go and used by every engineer on the platform monorepo.

Redaction is the load-bearing part

A development transcript is one of the most sensitive artifacts an engineering organization produces. It can contain whatever an engineer pastes into a terminal: API keys, customer identifiers, connection strings, personal data.

So the scrubbing worker sits between capture and the API call, and nothing reaches the model that has not been through it. This ordering is the whole design. A pipeline that sends first and filters later has already leaked; there is no recovering a request that has been made.

It is also why the pipeline is a worker rather than an inline hook. Capture has to be cheap and non-blocking, because a hook that slows down a developer’s own loop gets disabled inside a week, and a disabled hook produces nothing at all.

A finding nobody acts on is worse than no finding

The failure mode for this class of tool is not being wrong. It is being noisy. Engineers learn the shape of a bot that cries wolf within about three notifications, and after that it is furniture — still running, still costing money, read by nobody.

Two decisions kept it out of that category.

The model returns a typed structure, not prose. Each finding carries its category, its evidence and a confidence score, and the schema is what makes the next step possible at all. Free text cannot be thresholded, deduplicated or counted; a typed finding can be.

Only findings above a confidence threshold leave the pipeline. Everything else is discarded silently. This throws away real findings, and that is the correct trade: the cost of a missed finding is one unreported annoyance, and the cost of a false one is the credibility of every finding after it.

The output is a draft, and it is addressed to a person. It arrives in Slack pre-filled as a ticket, and a human accepts or rejects it. The pipeline never files anything itself, and it has no authority over anyone’s backlog — which is what made it welcome rather than resented.

Nobody holds the API key

Engineers do not hold the third-party credentials the tools need. A service holds them and brokers access, and engineers authenticate to that service with the OAuth 2.0 device authorization flow — the same flow a television uses, chosen for the same reason: the client is a terminal application with no browser and no safe place to keep a secret.

The instructions are part of the product

I also authored the agent instructions and skill definitions for three platform monorepos.

This turned out to matter more than it sounds. The quality of what an agent produces in a repository is bounded by how well the repository explains itself — where the tests live, what the conventions are, which commands to run and what each one is for. Writing those files is mostly the exercise of writing down what the team had only ever transmitted by tapping someone on the shoulder, which is worth doing whether or not a model ever reads it.