Edit this page


Zoat and Nina Rosa beside a multicolored network of understanding, agentic engineering, evidence and continuous software evolution.
A conceptual synthesis of understanding, bounded action, evidence and evolution—not an executed system or verified production outcome.

Contents

From understanding to accountable engineering

AI-Native Engineering is not merely about building software with AI. It is about transforming software engineering itself into an intent-driven, understanding-centered, and evidence-grounded process. Engineers increasingly coordinate agents, tools, and integrated workflows to investigate, plan, implement, test, and evolve systems while retaining responsibility for judgment, governance, and outcomes.

This changes the interface to engineering work: beyond editing code directly, the engineer can express intent, supply context, authorize bounded actions, inspect evidence, and revise understanding. The code remains essential, and direct programming remains appropriate whenever it is safer, simpler, or more precise. The shift is in how capabilities are orchestrated, not in handing accountability to a model.

The engineering environment can span ChatGPT, GitHub, IDEs, CI/CD, Skills, Hooks, MCP integrations, and SDKs, depending on the actual tools and permissions available. These are examples of potential components—not proof that a single platform supports every capability, nor a mandatory architecture. This standalone article develops the engineering practice without reproducing the symbolic narrative of The Developer’s AI Journey.

🧭 Reader question — role, process, or product?

Before choosing a label, ask what it describes: the developer’s role, the engineering workflow, or the capabilities of the software being built.

Predict before reading: If a team builds an LLM-powered support assistant using conventional pull requests, while another team uses coding agents to maintain an ordinary Rails app, which team is practicing “AI-Native Engineering”? The answer depends on what the phrase describes: the product or the way the product is engineered. We will separate these dimensions before the end.

For now, focus on a less fashionable and more useful question: What do we need to know before we allow an intelligent tool to change a working system?

We will follow a fictional Rails application that processes payment events in Sidekiq. A worker occasionally receives the same logical event twice. The business requirement is not “make the job run once”; it is “prevent duplicate financial effects while preserving legitimate retries and an audit trail.”

This is a deliberate choice. The safest solution to this problem might be deterministic code and a database constraint, not an LLM in production. AI can still help investigate, propose and review the change. The difference between using AI and depending on AI for correctness is the heart of this journey.

1. Engineering Judgment Beyond Prompts

The suggestion looks good. Would you ship it?

Imagine a coding assistant proposing the following change to a payment job:

# Illustrative pseudocode — not production-ready or executed.
def perform(event_id)
  return if PaymentEvent.exists?(external_id: event_id)
  PaymentEvent.create!(external_id: event_id)
  # Apply the payment effect...
end

It looks reasonable. But two workers could check exists? before either creates the record. Depending on the effect’s location and persistence guarantees, both could perform it. The method’s readability says little about its concurrency behavior.

Predict: Does a passing unit test with one invocation establish the “no duplicate effect” requirement? No. The test has not explored simultaneous calls, crash/retry windows, or partial external effects.

The engineering work is to specify the invariant, locate its authoritative enforcement boundary, identify failure paths and design evidence. Generated code is a candidate. Your review must answer whether the candidate protects the real business rule.

This matters beyond payment systems. AI-assisted development can make an initial solution easier to obtain without eliminating review cost. Research on AI and productivity is mixed and context-sensitive. For example, the METR study of experienced open-source developers in early 2025 examined a particular set of tasks, developers and tools; its findings should not be generalized to every developer or present-day model. DORA 2025 examines AI adoption in a broader organizational context, not the same experimental question. Verify exact numerical statements against their original methods before inserting them here.

Try this: Take a generated refactor from a test-backed Rails method. List one assumption about domain behavior, one untested failure path, and one observation that would change your assessment.

Checkpoint: Can you name a claim made by the suggested implementation and the evidence still missing to accept it?

2. Context, Architecture and Understanding

Why did the model miss the invariant?

“Fix duplicate payments” is an instruction. It is not yet sufficient context. An investigation needs to ask what counts as the same logical event, whether duplicate deliveries are expected, how events are persisted, when external calls occur, and what must remain observable when a retry fails.

Sidekiq’s official best-practices guidance explains why jobs should be idempotent: job execution is at least once, not an exactly-once guarantee. That property is easy to say but demanding to implement at the correct business boundary.

Write down the Use Case:

Field Hypothetical payment scenario
Actor / trigger Event delivery or job retry
Goal Apply one legitimate business effect per event
Exception flow Duplicate delivery, concurrent processing, timeout or crash
Constraint Audit the decision without discarding legitimate work
Evidence needed Persistence invariant, concurrency test, trace of side-effect boundaries

The Use Case describes a situation and its outcomes. Intent says what must hold. Inquiry uncovers unknowns. Engineering tests and implements bounded changes. Evidence supports specific claims; understanding may change as new observations arrive. These are complementary roles, not synonyms.

Context map — a model needs more than the job body

Mermaid diagram

Conceptual investigation map, not a diagram of executed tool calls. Files and logs provide candidate evidence; they do not certify a business invariant by themselves.

Try this: Compare the prompts “fix retry duplicates” and “inspect the job, database uniqueness rules, transaction boundary and tests in read-only mode; report unknowns before recommending changes.” What new questions become visible?

Checkpoint: Repository files and logs are useful sources. Which of their contents are actual observations, and which are interpretations not yet verified?

3. AI Assets and Bounded Agentic Workflows

It can run commands. Should we let it?

An assistant that drafts code and an agent that can invoke tools are not the same engineering capability. Permission to read, write, execute, publish or spend resources requires deliberate boundaries.

A useful mental model groups the supporting assets into three complementary lenses:

Understanding Execution Trust
Repository instructions, documented conventions, reusable Skills and domain context Agents, tool workflows, MCP connections and orchestration SDKs Permissions, Hooks where supported, tests, Evals, CI, auditability and observability
What does the assistant need to know? What bounded action may it perform? What has actually been checked, and who decides what happens next?
Three complementary lenses of understanding, bounded execution and trust, connected by contextual and evidence relationships.
Understanding, Execution and Trust are complementary lenses, not mandatory execution stages or proof that a system is safe.

These are editorial categories, not a formal vendor specification. MCP concerns integration and protocol contracts, not automatic authorization or correctness. GitHub Copilot agent mode documentation describes product capabilities; exact availability and controls depend on versions and settings and must be rechecked for the target host before publication.

For the Rails scenario, start narrow:

Inspect the worker, relevant models, migrations and specs. Do not edit files or run deployment commands. Identify the business idempotency key, possible duplicate side effects and unverified assumptions. Return a proposed change and tests for human review.

Once the inquiry is reviewed, a separate, explicitly authorized step may produce a diff. A test runner’s exit code is one observation, not evidence that every relevant case was covered.

🧰 Try this — permission boundary

Read: inspect jobs, migrations, specs and logs. Propose: report hypotheses and a reviewable plan. Write: only after approval. Release: a distinct human-authorized decision. The boundaries are examples to adapt to your actual environment.

Try this: Design one read-only task, one write-capable task and a separate human release gate. Which tool calls would be refused in each state?

Checkpoint: A tool call completed successfully. Which additional checks would establish whether the action was appropriate and the output reliable?

4. Evidence, Verification and Engineering Risk

The tests passed. What exactly do we know?

Rails transactions can make a group of database operations atomic within the transaction’s scope. They cannot reverse an already delivered HTTP request or other external effect simply because the database rolls back. See the Active Record transaction documentation.

Consider a proposed remedy: define a stable business idempotency key and enforce its uniqueness at the authoritative persistence boundary. Review how an external payment call, a database commit and the Sidekiq acknowledgment interact. An outbox-style design might help a particular integration, but it should be presented as a candidate architecture, not a solution established by this article.

🔎 Claim ≠ Evidence ≠ Understanding

Claim: “Retries cannot duplicate the business effect.”
Evidence required: database guarantees plus executed failure-seeking tests in the relevant environment.
Unknown: external service semantics, crash windows, concurrency assumptions.
Understanding: the provisional interpretation that can be revised when observations disagree.

A useful evidence ledger distinguishes:

Claim Possible evidence Still unknown
Two simultaneous jobs cannot create duplicate ledger entries A verified database constraint and a successfully executed concurrent test Real transaction isolation, production schema and external effects
Failed work can be retried safely Executed retry and recovery scenarios Every crash window and downstream response
The change is safe to release Reviewed PR, CI results and rollout plan Real production behavior after deployment

No tests have been executed for this illustrative case in this manuscript. Any future RSpec snippet should be run in a reproducible fixture and accompanied by the actual versions, commands and results.

Try this: Specify a test that starts two workers with the same event key. What should be asserted about database records and actual business effects? Which assertion would fail if the implementation used check-then-create without atomic protection?

Checkpoint: Identify one verified observation, one claim it supports and one risk that still requires investigation.

5. Evolving Existing Software Through Integrated Workflows

What changes when AI joins the engineering workflow?

Nothing about the Rails product itself needs to become “an AI application” for its engineering workflow to incorporate useful AI capabilities. The consequential change is the engineer’s role in directing the workflow: from mainly implementing a requested change by hand toward shaping the intent, orchestrating authorized capabilities, testing their results, and integrating what has been learned. The engineer may still write the critical code or reject automation altogether.

A practical environment might combine IDE or ChatGPT-based inquiry, GitHub repository and PR workflows, project instructions and Skills for conventions, MCP-connected tools for scoped system access, Hooks where supported for checks, SDKs for specialized orchestration, and CI/CD for reproducible verification. These elements are not inherently trusted or universally compatible. Permissions and evidence are separate from tool connectivity; each configuration must be reviewed in its real environment.

For the hypothetical payment-retry problem, a disciplined flow could connect:

  1. Intent: report the duplicate-effect risk and acceptance criteria.
  2. Inquiry: inspect repository state and relevant evidence with read-only tools.
  3. Plan: propose one bounded design change, negative tests and failure handling.
  4. Implementation: create a reviewable diff under approved permissions.
  5. Verification: run appropriate tests, review CI findings and record unknowns.
  6. Human decision: approve or reject the PR and a controlled release.
  7. Feedback: observe retry rates and unexpected effects, then revise the understanding.

This sequence is a proposed workflow, not a report of a completed deployment. The actions may form loops; a failed test or unexpected production observation should return the investigation to an earlier decision. NIST’s AI Risk Management Framework offers relevant Govern, Map, Measure and Manage perspectives, but it does not mandate the workflow above.

Proposed interaction sequence — not an executed deployment

Mermaid diagram

The sequence shows intended authority and evidence handoffs. It does not assert that any agent, CI job, or release ran for the hypothetical payment example.

Conceptual evolution of an existing codebase through inquiry, agent-assisted proposals, CI, human review and feedback.
A proposed evolution workflow for an existing codebase, not evidence of an executed deployment.

An agent could prepare the investigation and propose tests; CI could execute predetermined checks; the engineer still owns the decision about coverage, business impact and release authority. Integrating tools does not eliminate system boundaries or responsibility.

Try this: Sketch the developer–agent–repository–CI–reviewer sequence. Mark where read-only access becomes write access, where evidence is produced, and where a human must approve a consequential change.

Checkpoint: Which parts of this workflow are automated actions, and which are engineering judgments that cannot be inferred from an agent’s self-report?

6. AI-Native Products Versus AI-Native Engineering

Are we building an AI-Native product, or practicing AI-Native engineering?

It helps to ask two independent questions:

Question Example
What are we building? An application where model-backed reasoning is a core product capability
How are we building and evolving it? A conventional Rails service maintained through integrated, reviewable AI-assisted workflows

Two-axis thought experiment

  Conventional engineering workflow AI-integrated engineering workflow
Non-AI product Rails service developed conventionally Rails service evolved with bounded agents, tests and review
Model-powered product LLM feature built with ordinary development processes LLM feature built with integrated and governed agentic workflows

This matrix is an editorial distinction, not a maturity ranking or a standardized taxonomy.

These dimensions can intersect, but neither implies the other. A product may rely on an LLM while its development process remains conventional. A deterministic application may be built through agent-supported workflows without adopting any model in production.

AI-Native Engineering is our working editorial framing for a process that connects human intent, system context, bounded agentic work, verification, operational feedback and revisable understanding throughout software’s evolution. Public usage varies; it is not a single industry-wide standard or a badge awarded once an agent can run commands.

Nor does every problem benefit from an LLM. An explicit idempotency invariant is generally a deterministic engineering problem. Contrast it with ambiguous natural-language support-ticket classification: a probabilistic model might be useful there, but only after comparison with simpler baselines, error evaluation, human escalation and an operating plan. This second case is a thought experiment, not a measured result.

Try this: Classify these two axes for your team’s most recent feature. Which parts need intelligence, and which need straightforward, testable behavior?

Checkpoint: Can you explain why AI-enabled development does not automatically turn the resulting software into an AI-Native product?

Transfer — What would you do differently on Monday?

Choose one existing repository and a low-risk change. Write down its Use Case, desired outcome, known constraints, unknowns and source-of-truth files. Delegate only read-only inquiry first. Request a proposed diff and targeted tests. Record the evidence, note what remains unverified, and keep normal human review and deployment permissions.

Do this for a traditional software feature before choosing more complex multi-agent arrangements. The goal is not to install the largest toolkit: it is to make one engineering decision more informed and one change more verifiable.

Conclusion — Understanding in action

AI-Native Engineering, as framed here, is a change in the operating model of engineering, not merely a faster code-generation technique or a property of an AI-powered product. Intent defines what matters; inquiry and context establish what is understood and unknown; appropriately bounded tools and agents help carry out work; tests, reviews, and operational observations provide evidence; and the resulting understanding remains revisable.

The engineer increasingly acts as an orchestrator of engineering processes—connecting IDEs, ChatGPT, GitHub, Skills, Hooks, MCP, SDKs, and CI/CD where suitable—without surrendering direct coding skills, release authority, or accountability. Success means more than an agent completing a task: it means knowing why the change is warranted, what the evidence actually supports, what remains uncertain, and who is responsible for the next decision.

The approach applies both to new model-powered products and to established deterministic software, including systems already in production. The essential movement is from isolated prompts and implementations toward continuously informed, governed, and verifiable software evolution.

Next practical article (proposed): Building an AI-Native Engineering Environment with GitHub Copilot — From an Existing Rails Repository to Integrated Agentic Workflows. That practical tutorial can cover IDE setup, Instructions, Skills, MCP, permissions, PR/CI and release gates with executable examples. This article intentionally establishes the conceptual model first.

For a separate, symbolic exploration of the developer’s changing relationship with AI and evolving understanding, see The Developer’s AI Journey: From Prompts to AI-Native Engineering.

References and status

Editorial verification still pending: precise statistics if added, host-specific GitHub Copilot support, independent Sidekiq/RSpec test execution, claim-level reference review, Jekyll assembly and image QA. No claim in this manuscript should be mistaken for an executed lab result.