Starting Codex from zero with a real software project
The baseline for the Codex: Zero to Hero series: what we are building, what counts as evidence, and how we will judge the work.
This series is not a tour of Codex features.
I am more interested in a harder question: how do you make Codex useful inside normal software engineering work, where changes need scope, review, tests, constraints, and a trail of evidence?
That question is difficult to answer with toy examples. A small generated demo can show that an agent can write code, but it does not tell me much about how the workflow holds up after the second feature, the first awkward bug, the first bad abstraction, or the first time repository instructions start helping or getting in the way.
So this series starts with a real repository and a deliberately small application.
The public application repository is here:
https://github.com/rsantrod/codex-zero-hero-app
At the post-00 tag, it contains exactly one file:
.
└── README.md
No application code. No AGENTS.md. No CI/CD. No Docker configuration. No architecture documentation. No development tooling.
The application is the test bench
The application will become a personal finance app. That gives the series enough real engineering material to work with: accounts, transactions, categories, budgets, imports, reports, validation, persistence, UI state, and eventually operational concerns.
But the finance app is secondary.
The real subject is the engineering environment around Codex: how tasks are framed, how repository context is discovered, how changes are reviewed, how tests become part of the loop, when instructions help, when additional tooling is worth its maintenance cost, and where human judgment still has to lead.
The application will use fictional and sample financial data only. I do not want the series to depend on my own personal financial information, and I do not want readers to confuse a learning repository with a production finance system.
The initial technology direction is ordinary on purpose:
- Angular
- Spring Boot
- PostgreSQL
- Docker where useful
- Git
None of that is meant to be a claim that this is the best stack for every finance app. It is a mainstream enough stack to expose real frontend, backend, database, and workflow questions without making the series about novelty in the tools themselves.
The public repository is the experiment
The public repository, codex-zero-hero-app, is the reproducible engineering test bench. That is where the application and its Codex-facing engineering environment will evolve.
The series will follow a simple public evidence trail:
engineering question
-> Codex task
-> repository changes
-> evidence
-> human engineering review
-> conclusions
The repository state matters more than the intended narrative. If an experiment fails, produces an unimpressive result, or shows that some planned technique is not worth it yet, that is still useful material.
The conclusions in this series need to come from actual repository evidence and my engineering review: what changed, what ran, what failed, what needed correction, and whether I would be comfortable carrying the change forward.
The initial repository setup
I initialized the public repository using Codex Desktop against a local clone of the repository.
The prompt was intentionally narrow:
Initialize this repository as the application used in my "Codex: Zero to Hero" series.
The project will evolve into a personal finance application used as a real-world software engineering test bench while exploring how to work effectively with Codex.
For now, create only a README.md.
The README should briefly describe:
- the purpose of this repository;
- that the application will use fictional/sample financial data only;
- the initial technology direction:
- Angular
- Spring Boot
- PostgreSQL
- Docker where useful
- Git
- that both the application and its engineering environment will evolve incrementally throughout the series.
Keep the README concise. This is the initial state of the repository, not documentation for an application that already exists.
Do not scaffold the application yet.
Do not add AGENTS.md.
Do not add CI/CD.
Do not add Docker configuration.
Do not add architecture documentation.
Do not add development tooling.
Do not create any other files.
Before making changes, briefly tell me what you intend to do.
Then create only README.md.
Codex created exactly one file, README.md.

Codex Desktop after the initial repository setup: the constrained prompt and the generated README.md.
The generated README establishes the repository purpose, the fictional-data rule, the initial stack direction, and the fact that both the application and its engineering environment will evolve incrementally:
# Codex: Zero to Hero App
Initial technology direction:
- Angular
- Spring Boot
- PostgreSQL
- Docker where useful
- Git
Both the application and its engineering environment will evolve incrementally throughout the series.
That is enough for Post 00.
Not because the repository is useful yet. It is not. There is no application to run, no test suite to execute, no architecture to review, and no implementation quality to judge.
It is enough because the starting state is explicit and reproducible. Before asking Codex to build anything, the repository says what it is for, what kind of data it will use, and what stack direction the work will probably take.
Equally important, it says very little else.
Why start with almost nothing?
It would be easy to start by adding everything I expect to need later: repository instructions, architecture notes, a development environment, CI, Docker Compose, frontend and backend scaffolds, test commands, maybe even skills or MCP configuration.
Some of that may become useful. I expect at least some of it will. But adding it now would answer the wrong question.
I do not want to demonstrate a fully formed Codex workflow and then write articles explaining why each piece exists. I want to watch the workflow earn its shape.
So the first state of the public repository keeps those choices deliberately unresolved:
- No
AGENTS.mdyet. Repository instructions may become useful when actual engineering work gives us a sufficient reason to add them, but there is no evidence for that need yet. - No CI/CD yet. There is no code yet, and no evidence yet about what kind of validation loop this repository needs first.
- No Docker configuration yet. I do not want to add environment machinery before the first actual slice tells us what has to run.
- No architecture documentation yet. Architecture notes should follow real design pressure, not precede the first requirement.
- No additional agentic infrastructure yet. Skills, MCP, specialized tools, review automation, and similar mechanisms are all candidates. None of them has a job in this repository yet.
The evidence standard
The series will treat Codex work as engineering work, not as a magic trick.
Future posts should be grounded in evidence from the public repository:
- the exact task or prompt;
- the repository state before and after;
- files changed;
- relevant diffs;
- commands run;
- build and test output;
- failures and corrections;
- review notes;
- screenshots when the UI or workflow actually matters;
- tags, commits, or branches that let the state be found again.
That does not mean every post needs the same checklist. A small UI change and a database migration do not need the same evidence. A surface comparison and a debugging session should not be forced into the same shape.
But the rule is consistent: do not turn expectations into observations.
If Codex overreaches, that is evidence. If a structured task produces no meaningful improvement over a rough prompt, that is evidence. If a planned piece of process feels too heavy, that is evidence. If a feature works only after human correction, that correction belongs in the story.
The point is not to make Codex look good. The point is to understand where it is useful, where it is brittle, and what kind of engineering environment makes its work easier to trust.
What has not been tested yet
At this point, almost nothing has been tested.
I have not tested whether Codex can produce a good Angular and Spring Boot project structure from minimal guidance.
I have not tested whether the planned stack creates local setup friction.
I have not tested whether Codex needs repository instructions to stay within scope.
I have not tested whether Codex Desktop is the best surface for the first implementation task, or whether other Codex workflows will fit better for other tasks later.
I have not tested the application’s architecture, because there is no architecture yet.
That is the honest state of the project. Post 00 defines the experiment. It does not report implementation findings.
Where the series may go
The broad path is straightforward: start with a thin slice, review what Codex produces, improve the task framing, introduce persistent guidance when actual engineering work provides a sufficient reason for it, establish verification, and then move into more realistic engineering problems.
Later, the series may explore debugging, review workflows, data imports, external integrations, background processing, reusable instructions, skills, MCP, delegated work, larger refactors, and team adoption.
But those are directions, not promised conclusions. The public repository has to earn each step one experiment at a time.
The next step is Post 01: ask Codex to build the first thin vertical slice of the finance application with minimal repository guidance.
That should give us the first useful baseline: what Codex infers correctly, where it reaches too far, what it misses, and what a human engineer has to review before the work can be trusted.