Production AI systems · Berlin

Build AI systems that improve from evidence

I design production AI systems across software engineering, agents, evaluation, retrieval, security, and infrastructure—turning traces and failures into controlled experiments, reliable releases, and safer behavior.

Explore the work Run the labs

Byte, the Edge of Context beaver

The production problem

The demo is not the hard part.

Production agents fail in the seams between memory, retrieval, tools, permissions, evaluation, and release control. Improving one score is not progress if the change weakens reliability, security, or cost.

Three connected engineering pillars

One governed system
01

Closed-Loop AI Engineering

Turn traces and failures into bounded experiments, independent evaluation, and evidence-backed promotion decisions.

02

Agent Assurance & Security

Test what an agent can be manipulated into doing before it ships—and keep policy outside the optimizer.

03

Production Agent Systems

Design memory, retrieval, tools, harnesses, runtimes, and observability as one governed system.

The operating loop

Automate the experiment, not the authority to declare it safe.

  1. Observe
  2. Diagnose
  3. Propose
  4. Experiment
  5. Evaluate
  6. Challenge
  7. Promote
  8. Learn

Systems, experiments, and evidence

See all work →

Agent systems

Engineering the Agentic Stack

A six-part architecture covering reasoning, memory, tools, security, runtime, and the harness around the model.

6 parts · reference system

Evaluation

AI agent evaluation in production

Turn traces into regression suites and connect them to a runnable RAG evaluation harness with bounded claims.

Article + lab · traces to tests

Retrieval & language

Search, OCR, and NER

Field guides for ranking, document understanding, structured extraction, and the systems that connect them.

Field guides · runnable examples

About

The systems around the model are the product.

I’m Slava Dubrov, a Staff AI System Engineer with a doctoral degree in AI diagnostics. I work across software engineering, agent systems, evaluation, retrieval, security, and production infrastructure.

More about my work →

Explore further

Follow the evidence behind the ideas.

Read the engineering work, run the public artifacts, or continue through the writing behind the systems.

Explore the work Read Edge of Context