Senior systems and data engineer working on AI evaluation, model understanding, and reliable AI systems.

I have spent 25+ years building production software, data, analytics, and enterprise systems. My current work investigates how AI systems represent their situation, why they behave as they do, and how those behaviors can be evaluated against evidence and constraints.

Governed Natural Language to SQL Generation: A Behavioral and Mechanistic Analysis

Investigated how language models encode task and constraint information, combining activation analysis with behavioral evaluation across governed NL-to-SQL settings.

World-state reasoning in intelligent systems

Exploring when AI systems need explicit representations of changing state, how they update those representations, and how failures can be evaluated against evidence and constraints.

  1. 01

    When does successful model behavior conceal an incorrect or inconsistent understanding of its situation?

  2. 02

    How can we determine why an AI system took a particular action, rather than merely whether the action was correct?

  3. 03

    Which parts of an intelligent system’s world state should be learned, inferred from context, retrieved through tools, or represented explicitly?

My background spans software engineering, data systems, distributed systems, analytics, geospatial computing, and enterprise architecture. I am now using that experience to develop a more empirical research program around AI systems, evaluation, and model behavior.