Codex vs Claude Code Skills: AI Coding Is Moving Beyond Prompts
For engineering leaders, the opportunity may therefore be less about choosing a permanent winner between OpenAI and Anthropic and more about asking a different question:

Codex vs Claude Code Skills: What They Mean for Software Teams
The discussion around AI coding tools has mostly been about models: which one produces better code, understands a repository faster, or handles a difficult debugging problem more reliably.
That is still relevant, but I think another development deserves more attention.
Both OpenAI Codex and Anthropic Claude Code are moving toward reusable Skills: packaged instructions, references, scripts and workflows that tell an AI coding agent how a particular type of work should be carried out.
That changes the conversation considerably.
The useful part of an experienced developer is not only their ability to write code. It is also everything they have learned around the code: how the company structures APIs, how deployments are handled, which shortcuts tend to cause problems later, how integrations should fail safely, what needs to be reviewed before a release, and which conventions matter in production.
Until now, most of that knowledge has lived in documentation, repositories and people's heads.
Skills offer a practical way of putting more of it directly into the workflow of an AI agent.
What a Skill actually is
At a basic level, a Skill is usually built around a SKILL.md file.
The file describes what the Skill is for and how the agent should perform the task. It can also sit alongside reference documents, scripts, examples or other supporting material.
A simple repository might contain something like:
skills/
api-review/
SKILL.md
references/
api-standards.md
security-checklist.md
scripts/
validate-api.sh
The value here is fairly straightforward.
Without a Skill, somebody might repeatedly tell an AI coding assistant:
- check authentication
- validate request handling
- review error responses
- confirm API versioning
- check logging
- run the relevant tests
With a well-written Skill, those expectations already exist in a reusable form.
The developer can ask for an API review without having to explain the organisation's review process from the beginning every time.
OpenAI describes Skills as reusable instructions and supporting files that can represent anything from company conventions to multi-step workflows. Anthropic's implementation follows a similar idea through SKILL.md files that Claude can load when the task is relevant.
This is closer to reusable engineering procedure than prompt engineering.
Codex and Claude Code are heading in a similar direction
The implementation details are not identical, but the overall design is surprisingly similar.
Both systems allow the agent to know that a Skill exists without permanently loading every instruction and reference document into context.
When a task calls for that Skill, the relevant material can be brought in.
That matters more than it may initially sound.
A mature engineering team could easily have Skills covering areas such as:
dotnet-api-development
code-review
security-review
azure-deployment
database-migration
integration-testing
incident-analysis
It would make little sense to load all of those instructions every time somebody asked the agent to rename a method or fix a unit test.
Instead, the agent can discover the appropriate capability as needed.
That keeps context focused and makes a large library of organisational guidance more practical.
Where Codex is particularly interesting
Codex also has the concept of AGENTS.md, which works well for persistent repository guidance.
For example:
Use .NET 10.
Public APIs must not expose supplier-specific models.
Use structured logging for external integrations.
All external service calls require timeout handling.
Follow the repository's dependency injection conventions.
These are not really individual tasks. They describe how work in that repository should normally be done.
Skills can then deal with specific activities.
For example:
AGENTS.md
Repository architecture
Coding conventions
Platform constraints
General engineering rules
Skills
create-api
review-api
add-provider
database-migration
production-readiness
I like this separation because it resembles how engineering teams already operate.
There are standards that apply almost everywhere, and there are separate procedures for particular jobs.
Codex can use both.
Claude Code takes a slightly different route
Claude Code has made Skills feel quite natural inside its command-line workflow.
A Skill can be selected automatically based on the request, while developers can also invoke one directly with a command such as:
/code-review
Anthropic has also extended the Skill model with controls around invocation, tool access, subagents and dynamic context.
Dynamic context is particularly useful.
A Skill can gather the current repository state before the model performs its analysis. A review Skill, for example, could obtain the current Git diff and then apply the organisation's review checklist to those actual changes.
This turns the Skill into more than a document.
It becomes a small workflow combining instructions, live context, tools and model reasoning.
That is probably where this concept becomes most valuable.
The real comparison is not Codex versus Claude
It is tempting to turn this into another product comparison and decide which implementation is better.
There are differences worth considering.
Claude Code has a very natural command-oriented experience around Skills and provides useful controls around invocation and context.
Codex fits Skills into a broader setup that includes repository instructions, tools and agent workflows.
But I suspect most serious engineering organisations will eventually use more than one coding model.
The more important question is therefore:
Can the engineering knowledge survive when the model changes?
This is where a standardised Skill format becomes interesting.
If a company keeps its API review process, deployment conventions and integration standards in reusable Skill directories, those assets do not necessarily have to belong permanently to one AI vendor.
A structure might look like this:
company-skills/
api-review/
SKILL.md
references/
scripts/
security-review/
SKILL.md
provider-integration/
SKILL.md
references/
deployment/
SKILL.md
scripts/
The model becomes the reasoning layer.
The organisation continues to own the process.
For businesses that expect AI tooling to change quickly over the next few years, that distinction is important.
This matters even more in domain-heavy software
The idea becomes particularly useful in industries where the difficult part of development is not syntax.
Travel technology is a good example.
An experienced travel-tech team accumulates knowledge about things such as:
- supplier authentication behaviour
- fare and offer normalisation
- GDS-specific error patterns
- booking and ticketing flows
- retries and idempotency
- hotel cancellation policies
- provider certification
- PNR handling
- reconciliation
- production support
A generic coding model knows software.
It does not automatically know how one particular company's travel platform has learned to deal with Sabre, Amadeus, Expedia, NDC providers or other suppliers over several years.
That experience could gradually be captured in Skills such as:
air-provider-integration
hotel-provider-onboarding
fare-normalisation
pnr-diagnostics
supplier-certification
booking-failure-analysis
The Skill would not replace the architect or the domain expert.
It would make their knowledge available much more consistently across everyday development.
That could be especially useful when new engineers join a project.
Instead of expecting them to discover years of decisions through old tickets, documentation and code review comments, the coding agent working with them can apply some of those standards from the beginning.
There is also a governance problem
Skills sound attractive because they make good practices repeatable.
The opposite is also true.
They make bad practices repeatable.
Suppose a Skill tells an agent to retry every failed request five times.
That instruction may look harmless until it is applied to an operation that is not idempotent.
A badly designed Skill could spread an architectural mistake across multiple services far faster than a developer making the same mistake manually.
For that reason, organisations will probably need to treat important Skills more like source code than documentation.
They should have owners.
Changes should be reviewed.
Security-sensitive Skills should be controlled.
Important workflows should be tested against real scenarios.
Old instructions should be updated when the platform changes.
OpenAI's documentation also points out the security implications of Skills, particularly when instructions, scripts and external tools are combined. Anthropic provides controls around who can invoke Skills and which tools they are allowed to use.
As these systems become more capable, Skill governance may become a normal part of engineering governance.
A useful way to think about the architecture
I currently see the pieces like this:
Model
Reasoning
Repository
Application context
AGENTS.md / CLAUDE.md
Persistent project guidance
Skills
Repeatable engineering procedures
Tools
Actions and external systems
Engineer
Judgement and accountability
None of those elements is particularly revolutionary on its own.
Together, however, they create something quite different from the coding assistants we were using a few years ago.
The agent is no longer simply generating code from a request.
It can increasingly work within an operating model defined by the engineering organisation.
Skills may become part of a company's technical IP
Companies usually think about technical intellectual property in terms of source code, data, algorithms and architecture.
AI agents may add another category.
A well-developed library of Skills can contain the practical experience accumulated from hundreds of deployments, incidents, integrations and architectural decisions.
That library might eventually describe how the company:
- creates services
- integrates suppliers
- reviews security
- diagnoses incidents
- migrates databases
- tests integrations
- prepares releases
- handles production failures
Two companies could use the same underlying AI model and still get very different results because one has spent years developing better operational knowledge around that model.
That feels like a more durable advantage than simply having access to whichever coding model happens to lead the benchmark table this month.
Where I think this is going
Software development has always introduced new ways of making knowledge reusable.
Functions made code reusable.
Libraries made capabilities reusable.
APIs made systems reusable.
Infrastructure as Code made infrastructure definitions reusable.
CI/CD pipelines made delivery processes repeatable.
Skills could do something similar for AI-assisted engineering work.
They provide a place to describe not merely what code should be generated, but how engineering work should be performed.
That distinction is important.
The next stage of AI development is unlikely to be defined only by models becoming better programmers.
It will also depend on how effectively organisations can give those models the right context, processes and domain knowledge.
Codex and Claude Code are both moving in that direction.
For engineering leaders, the opportunity may therefore be less about choosing a permanent winner between OpenAI and Anthropic and more about asking a different question:
Which parts of our engineering knowledge are valuable enough to capture as reusable capabilities?
That is where Skills become more than another feature in an AI coding tool.
They become part of how the engineering organisation itself operates.
Have a similar workflow?
Discuss what a focused first project could involve. Tell us which article you read.