The next problem is not generating more code. It is maintaining semantic continuity from human intent to production evidence across a lifecycle increasingly executed by agents.
AI is moving upward toward human intent.
Increasingly, the starting point is not a file, a function, or even a programming language. It is a description of what should exist, what should change, what constraints matter, and what outcome is expected.
Agents can translate that intent into architecture, code, tests, configuration, deployment changes, and operational actions.
But that shift creates a less obvious problem underneath the interface:
Meaning still has to survive the translation.
A requirement must remain connected to the evidence that justified it. An architectural decision must remain connected to the trade-offs behind it. Code must remain connected to the intent it implements. Tests must remain connected to the behavior they verify. Deployment and runtime evidence must remain connected to the changes that produced them.
Once agents begin participating across that entire lifecycle, source code alone is no longer enough context.
The engineering system needs a semantic layer that preserves:
why something exists
what it means
what constrains it
what implements it
what verifies it
what happened when it ran
what we learned afterward
That leads to the question behind this post:
What semantic infrastructure must exist underneath AI so agents can understand, verify, and operate software reliably?
Contents
- The abstraction moved up
- Below intent, semantics still matter
- The SDLC is the real test
- Semantic continuity
- From source code to a program database
- Rubyβs paradox: expressiveness and inspectability
- A DSL should compile into behavior and understanding
- One declaration, many projections
- From program database to engineering semantic graph
- Runtime observability becomes an agent interface
- GitHub Copilot across the lifecycle
- What should an agentic engineering platform expose?
- The opportunity for Ruby
- Conclusion
- References
The abstraction moved up
The history of programming is a history of moving the interface upward.
Machine code exposed the machine. Assembly gave names to instructions. Higher-level languages let us describe algorithms and data structures. Frameworks let us speak in the vocabulary of domains.
AI raises that boundary again.
The addition in the middle matters.
If AI is the human-facing abstraction, the system still needs a representation that preserves more than the latest prompt. It needs to preserve meaning: why something exists, which constraints apply, which evidence supports it, how it is implemented, what can change it, and what happened when it ran.
Without that layer, AI may generate implementation faster while the engineering system becomes harder to understand as a whole.
Below intent, semantics still matter
It is tempting to conclude that programming languages matter less once agents write more code.
That is only partly true.
Valimβs argument is more subtle. If agents are less sensitive to boilerplate and typing ergonomics, then optimizing a language primarily for the act of typing may become less strategically important. But languages still encode semantics, constraints, execution models, locality, and guarantees. Those properties may become more important, not less, when humans no longer inspect every generated line.
The design question changes from:
Which language is nicest for a human to type?
to:
Which language and runtime expose the strongest combination of meaning, guarantees, inspectability, and observable execution?
That is a much richer question.
Ruby, Elixir, Rust, Java, and other languages are not interchangeable token targets. They expose different semantic worlds.
The AI layer may make it cheaper to move between those worlds. It does not erase the consequences of choosing one.
The SDLC is the real test
A coding agent can survive with repository context.
An engineering agent cannot.
Once AI participates across discovery, planning, architecture, implementation, testing, review, release, production, and feedback, source code becomes only one slice of the context it needs.
Consider the information required at each phase:
Discovery β customer evidence, product intent, constraints
Planning β priorities, acceptance criteria, dependencies
Architecture β ADRs, boundaries, contracts, trade-offs
Development β repository semantics, instructions, implementation context
Testing β expected behavior, invariants, edge cases
Review β intent, policy, security rules, diff, architecture
Release β gates, deployment state, approvals, provenance
Production β traces, state, incidents, dependencies, runbooks
Feedback β outcomes, failures, customer signals, learning
If every agent reconstructs this independently, the system develops semantic drift.
The architecture agent may understand one version of the intent. The coding agent another. The reviewer sees only the diff. The incident agent sees only the logs.
The real problem is not context size.
It is continuity of meaning.
Semantic continuity
The useful unit is not a prompt.
It is a semantic thread.
There is a useful philosophical echo here: Krishnamurti repeatedly challenged the tendency to confuse accumulated knowledge with direct perception. In engineering terms, that is a reminder that a semantic model can support understanding without becoming identical to the reality it represents.
The engineering system should be able to answer questions that cross these boundaries:
Which requirement caused this code path to exist?
Which ADR justified this dependency?
Which tests verify the policy behind this behavior?
Which release introduced the change?
Which incident produced the feedback that changed the requirement?
Those are not code-completion questions.
They are system-understanding questions.
A model can connect fragments, but the map should never be confused with the whole.
From source code to a program database
Valim proposes an important shift for agentic tooling.
Traditional IDE interfaces are shaped around how humans browse software: files, lines, columns, symbols, βgo to definition.β Language servers already know much more than the interface exposes: references, call graphs, types, and sometimes data-flow relationships.
Why make agents hop through files one fragment at a time?
A program database could expose the program as something queryable.
For example:
Find every public function that can eventually reach this operation.
Find all execution paths where this value can become nil.
Find all callers that cross this trust boundary.
Find every component whose output eventually reaches this deployment action.
The important shift is:
source files
β
parsed structure
β
semantic relationships
β
queryable program model
For humans, βgo to definitionβ is convenient.
For agents, composing structural queries may be more powerful.
Rubyβs paradox: expressiveness and inspectability
Ruby is especially interesting because the same features that make it expressive can make structural reasoning harder.
Rubyβs history is full of mechanisms that let libraries create beautiful internal languages:
has_many :orders
before_action :authenticate_user!
task deploy: :environment
expect(result).to be_success
That expressiveness is a strength.
But highly dynamic behavior can also create action at a distance:
monkey patching
implicit hooks
runtime rebinding
method_missing
open classes
dynamic dispatch
Valim explicitly calls out locality as important for agentic tooling. If behavior can be changed from far away, even a good program database has more work to do.
So the opportunity for Ruby is not simply:
more metaprogramming.
It may be:
expressive metaprogramming backed by an explicit semantic model.
Ruby can remain delightful to read while exposing enough structure for agents to reason about what the DSL actually means.
A DSL should compile into behavior and understanding
This is where the previous postβs idea can evolve.
Consider:
agent :architect do
knows :architecture
uses :github
produces :adr
end
The naive interpretation is:
DSL
β
runtime behavior
A more useful AI-native interpretation is:
DSL
βββ behavior
βββ understanding
That gives us a principle worth keeping:
A successful AI-native DSL should compile into both behavior and understanding.
One declaration, many projections
This also changes the meaning of one declaration, many projections.
In the previous post, one Ruby declaration could generate JSON, Markdown, Mermaid, agent profiles, skills, and runtime configuration.
Now add one more target:
a queryable semantic representation of the engineering system itself.
A declaration such as:
agent :discovery do
goal "Understand the problem"
knows :product
uses :github
produces :discovery_report
policy do
distinguish :evidence, :inference, :unknowns
end
end
could project into:
Agent profile
Skill package
MCP bindings
Documentation
Mermaid diagrams
Runtime configuration
Policy declarations
Semantic graph
Program queries
Provenance links
The Ruby source is no longer merely configuration.
It becomes a semantic authoring surface.
From program database to engineering semantic graph
A program database is necessary, but an agent that spans the SDLC needs something larger.
It must connect program structure to the history and intent around the program.
The graph might connect:
Requirement
β
ADR
β
Code
β
Tests
β
Policy
β
Deploy
β
Incident
β
Feedback
But the useful representation is not merely chronological.
It should preserve relationships such as:
Requirement R-17
βββ justified by interview evidence E-3
βββ shaped by ADR-12
βββ implemented by component C-8
βββ verified by tests T-4 and T-5
βββ constrained by policy P-2
βββ deployed in release 142
βββ implicated in incident INC-42
βββ revised after feedback F-9
That is closer to an engineering semantic graph than a source-code database.
A semantic graph can improve legibility without claiming completeness. Knowledge helps us see relationships; it does not exhaust reality.
It gives different agents different projections of the same underlying meaning.
Runtime observability becomes an agent interface
Debuggers are interfaces designed around human attention.
We set a breakpoint, stop the world, inspect variables, step, inspect again.
Agents do not have the same limitation.
They can collect traces, correlate signals, issue structural queries, and compare state across many processes or requests.
That suggests a different future interface:
runtime state
β
query
β
correlation
β
hypothesis
β
verification
Valim points to the BEAM as an example of a runtime already rich in introspectable concepts: processes, supervisors, sockets, ETS tables, applications, and message queues.
That does not make Elixir automatically βthe AI language.β
It demonstrates a property that may become strategically important:
runtime state as a first-class, queryable interface for agents.
The same question applies to Ruby, Rails, JVM systems, containers, Kubernetes, queues, databases, and observability platforms.
How much of the running system can an agent inspect safely and structurally rather than infer from a pile of logs?
What we know, what is happening, and what we think it means should remain distinct.
That distinction matters especially in agentic systems, where retrieved knowledge, live evidence, and model inference can easily collapse into one another if the architecture does not keep them explicit.
GitHub Copilot across the lifecycle
This is where the discussion becomes concrete.
GitHub Copilot is no longer a single autocomplete surface.
GitHub documents Agent Skills as reusable task-specific packages that can work with the Copilot cloud agent, code review, Copilot CLI, the GitHub Copilot app, and agent mode in Visual Studio Code. Copilot code review can also use relevant Skills and MCP servers to bring repository and external context into a review.
Copilot CLI exposes another set of primitives: custom agents, MCP servers, hooks, skills, custom instructions, and persistent repository memory.
Those mechanisms are powerful, but they also reveal the architectural problem.
Each surface needs a different slice of meaning:
Discovery agent β product evidence
Planning agent β priorities and constraints
Architecture agent β system boundaries and ADRs
Coding agent β repository semantics
Review agent β intent + policy + diff + architecture
Release agent β gates + provenance
Operations agent β runtime state + runbooks + incidents
If every surface receives an independently assembled context packet, the system worksβbut the same semantics are encoded repeatedly.
A shared semantic model changes that.
Engineering Semantic Model
β
βββββββββββββββββββββΌββββββββββββββββββββ
βΌ βΌ βΌ
Product view Program view Runtime view
β β β
discovery/planning coding/review release/operations
β β β
βββββββββββββββββββββΌββββββββββββββββββββ
βΌ
Different Copilot surfaces
The AI asset is no longer the source of truth.
The AI asset becomes a projection of the source of truth.
That distinction is important.
What should an agentic engineering platform expose?
The pieces now separate into clearer layers.
SEMANTICS
product intent Β· architecture Β· program model Β· lifecycle relationships
KNOWLEDGE
RAG Β· docs Β· ADRs Β· runbooks Β· history
CAPABILITIES
Skills Β· custom agents
ACCESS
MCP Β· tools Β· APIs
CONTROL
instructions Β· policies Β· hooks Β· approvals
EXECUTION SURFACES
IDE agent mode Β· Copilot CLI Β· Copilot app/cloud agent Β· code review
EVIDENCE
tests Β· traces Β· deployments Β· incidents Β· feedback
The mistake would be to call all of these βAI assets.β
They are not equivalent.
A Skill tells an agent how to perform a task.
MCP gives it access to tools and external systems.
RAG retrieves knowledge.
A program database exposes program structure.
Runtime observability exposes live execution state.
An engineering semantic graph connects those layers to why the system exists and how it evolved.
That is the deeper infrastructure.
The opportunity for Ruby
This brings us back to Ruby.
Ruby does not need to win because agents find its syntax easy to generate.
That would be a weak and temporary advantage.
Its more interesting opportunity is architectural.
Ruby has an unusually rich tradition of turning complicated machinery into domain language:
has_many :orders
describe Order do
it { is_expected.to be_valid }
end
task deploy: :environment
The AI era creates a new domain vocabulary:
Agent
Skill
Knowledge
Tool
Policy
Evidence
Handoff
Flow
Runtime
Incident
Feedback
Ruby can express that vocabulary elegantly.
But the next step is crucial.
The DSL should not hide the system behind magic. It should create an explicit model that both humans and agents can inspect.
Something like:
engineering_system :caseflow do
requirement :sla_escalation do
evidence :customer_interviews, :support_ticket_4312
constrained_by :notification_policy
end
architecture :notifications do
decided_by :adr_003
implements :sla_escalation
end
verification :sla_request_specs do
verifies :sla_escalation
end
production :caseflow_api do
exposes :traces, :deployments, :incidents
end
end
That could generate behavior.
It could also answer:
Why does this feature exist?
Which evidence justified it?
Which ADR shaped it?
Which tests verify it?
Which policy constrains it?
Which deployment introduced it?
Which incidents changed our understanding of it?
That is the possibility I find most interesting:
Ruby as a semantic compiler for software engineering intent.
Not merely a runtime language.
Not merely a pleasant DSL.
A language that turns intent into both executable behavior and queryable understanding.
Conclusion
The abstraction is moving upward.
Humans increasingly describe intent.
Agents increasingly translate that intent into implementation.
But that movement creates a new responsibility underneath AI.
We need systems that preserve what the implementation means.
JosΓ© Valimβs program database and runtime-observability ideas point toward better interfaces for coding agents.
The SDLC expands the problem further: agents need product intent, architecture, code semantics, policy, verification evidence, deployment history, runtime state, incidents, and feedback to remain connected.
GitHub Copilotβs expanding surfaces make this problem visible in practice. Skills, custom agents, MCP, hooks, code review, CLI, and cloud execution can all benefit from shared semantics rather than isolated context reconstruction.
And Ruby gives us an intriguing mechanism for expressing that shared model.
As AI moves closer to human intent, software engineering needs an equally strong semantic layer underneath it.
That layer must preserve why a system exists, how it is structured, what constrains it, what verifies it, and what happens when it runs.
Programming languages, program databases, runtime observability, engineering knowledge, and agent infrastructure may increasingly become different projections of that same semantic system.
The future language-design question may therefore be neither:
Which language is fastest?
nor:
Which language is easiest for AI to write?
It may be:
Which language and engineering platform let humans and agents understand, verify, query, and observe the systems they create together?
The challenge is no longer simply to make agents capable of producing software.
It is to make software itself legible to the humans and agents responsible for it.
And perhaps the most important requirement for an AI-native engineering system is surprisingly simple:
It should produce both behavior and understanding.
Support material
The companion package contains the Mermaid sources, the concept map, example semantic queries, and a small Ruby sketch of an engineering semantic model.
References
- GitHub Docs β About GitHub Copilot code review
- GitHub Docs β Adding agent skills for GitHub Copilot
- GitHub Docs β About GitHub Copilot CLI
- GitHub Docs β GitHub Copilot CLI command reference
- GitHub Docs β Using MCP servers with the GitHub Copilot SDK
- Previous post β Pencils Down, Intent Up
- Previous post β Ruby in the Age of AI
- Previous post β AI Across the SDLC with GitHub Copilot
- JosΓ© Valim β Evolving programming languages in the AI era