Asking Codex for the first feature before the repository knows anything

What happened when Codex was asked to build the first Angular, Spring Boot, and PostgreSQL slice from an almost empty repository — and the small defect that survived review.

Post 00 left the public application repository in a deliberately sparse state.

At the post-00 tag, the repository contained one file:

.
└── README.md

No Angular scaffold. No Spring Boot scaffold. No Docker Compose file. No test setup. No AGENTS.md. No architecture notes. No repository-specific instructions for Codex.

That was the point. Before improving the engineering environment around Codex, I wanted a baseline: what happens when the first real implementation request lands in a repository that only says what it is for and what stack it expects to use?

This post is that baseline.

The public repository evidence is available directly:

  • Starting state: post-00, commit be65789d999d22bcec43d1ac3310104f3f4446d8
  • Final baseline: post-01, commit 3dbbde080b91cfc3521d2169e3a665b0ad6bc52d
  • Repository diff: post-00...post-01

The application is still a test bench, not a production finance product. The data is fictional. The engineering work, repository changes, run behavior, review, and defects are real.

The request

The task was intentionally realistic but bounded. I did not ask for a complete finance application. I asked for the first slice: accounts, recent transactions, persistence, a read-only API, and a simple Angular screen.

This is the prompt I gave Codex Desktop in the public application repository:

Build the first thin vertical slice of this personal finance application.

Starting point: the repository currently only has the README from the initial series setup. There is no application scaffold yet.

Please create a minimal full-stack application using:

- Angular for the frontend
- Spring Boot for the backend
- PostgreSQL for persistence, with fictional/sample data only
- Git-friendly project structure and documentation

For this first slice, keep the product scope small:

- accounts have at least an id, name, type, and current balance
- transactions have at least an id, account, date, description, amount, and category
- the backend exposes read-only REST endpoints for accounts and recent transactions
- the frontend shows a simple page with the accounts and recent transactions loaded from the backend
- include a small amount of fictional seed/sample data so the first screen can show something useful

Please keep this as a first working slice, not a complete finance app.

Do not add authentication.
Do not add budgets.
Do not add CSV imports.
Do not add reporting or analytics.
Do not add CI/CD.
Do not add AGENTS.md or Codex-specific instructions.

If Docker or Docker Compose is useful for running PostgreSQL locally, add only the minimal configuration needed for this slice.

Before editing, briefly inspect the repository and tell me your implementation plan.
Then implement the slice, run the relevant build/test checks you can run locally, and update the README with concise instructions for running the app.

There are two things I care about in that prompt.

First, the scope is small enough to review. It forces Codex to make real choices across frontend, backend, API, persistence, local runtime, verification, and documentation, but it avoids product features that would drown the baseline in unrelated complexity.

Second, the prompt includes normal engineering constraints but not hidden experiment intent. It tells Codex what not to build, because that is part of the request. It does not tell Codex that this is a test of minimal guidance, or that later posts may introduce repository instructions, or what conclusions I expect.

What changed

Codex produced a conventional small full-stack repository:

.
├── README.md
├── docker-compose.yml
├── backend/
│   ├── pom.xml
│   └── src/
└── frontend/
    ├── angular.json
    ├── package.json
    └── src/

The post-00...post-01 comparison changed 33 files. Most of that is normal scaffold and lockfile volume: Maven wrapper scripts, an Angular lockfile, Spring Boot source, Angular source, configuration, and tests.

The implementation stayed inside the requested product scope. It did not add authentication, budgets, imports, reporting, analytics, CI/CD, or AGENTS.md.

The README was updated with the expected local run path:

docker compose up -d postgres
cd backend
./mvnw spring-boot:run
cd frontend
npm install
npm start

That is not a sophisticated developer environment. It is enough for the first slice.

Backend shape

The backend is a small Spring Boot application using JPA, Flyway, PostgreSQL, and read-only REST endpoints.

The database schema is explicit Flyway SQL. It creates accounts and transactions, adds a recent-transaction index, and inserts fictional seed data:

create table accounts (
    id bigint primary key,
    name varchar(120) not null,
    type varchar(40) not null,
    current_balance numeric(12, 2) not null
);

create table transactions (
    id bigint primary key,
    account_id bigint not null references accounts(id),
    transaction_date date not null,
    description varchar(180) not null,
    amount numeric(12, 2) not null,
    category varchar(80) not null
);

create index idx_transactions_recent on transactions (transaction_date desc, id desc);

The application validates the schema rather than letting Hibernate create it implicitly:

spring:
  datasource:
    url: ${SPRING_DATASOURCE_URL:jdbc:postgresql://localhost:5432/personal_finance}
    username: ${SPRING_DATASOURCE_USERNAME:finance_app}
    password: ${SPRING_DATASOURCE_PASSWORD:finance_app}
  jpa:
    hibernate:
      ddl-auto: validate
    open-in-view: false
  flyway:
    locations: classpath:db/migration

That is a reasonable first-slice choice. The schema is reviewable, the seed data is visible, and the app will fail if the entity mapping and database drift apart.

The controllers return response records instead of serializing JPA entities directly. That keeps the first API contract separate from the persistence model without adding much ceremony: AccountResponse and TransactionResponse are small DTO records mapped from the JPA entities at the controller boundary.

The recent transactions endpoint also has a useful bit of defensive shape: the caller can request a limit, but the controller clamps it between 1 and 50 and sorts deterministically:

@GetMapping("/recent")
public List<TransactionResponse> recentTransactions(
    @RequestParam(defaultValue = "10") int limit
) {
    int safeLimit = Math.max(1, Math.min(limit, 50));
    PageRequest page = PageRequest.of(
        0,
        safeLimit,
        Sort.by(Sort.Direction.DESC, "date").and(Sort.by(Sort.Direction.DESC, "id"))
    );

    return transactionRepository.findAll(page)
        .stream()
        .map(TransactionResponse::from)
        .toList();
}

The repository uses an entity graph for the account relationship needed by the transaction response:

public interface TransactionRepository extends JpaRepository<FinancialTransaction, Long> {

    @Override
    @EntityGraph(attributePaths = "account")
    Page<FinancialTransaction> findAll(Pageable pageable);
}

That matters because spring.jpa.open-in-view is disabled. Without either an explicit fetch strategy or a transaction boundary that covers mapping, this kind of response mapping can easily run into lazy-loading problems. For a first slice, the chosen solution is simple and good enough.

Frontend shape

The frontend is a compact Angular app. For this baseline, Codex kept everything in main.ts: API calls, local state, derived total balance, loading/error state, and the template.

The data loading is direct:

constructor() {
  Promise.all([
    this.loadAccounts(),
    this.loadTransactions()
  ])
    .catch(() => {
      this.error.set('Could not load the finance snapshot. Check that the backend is running on port 8080.');
    })
    .finally(() => this.loading.set(false));
}

private loadAccounts(): Promise<void> {
  return new Promise((resolve, reject) => {
    this.http.get<Account[]>('/api/accounts').subscribe({
      next: accounts => {
        this.accounts.set(accounts);
        resolve();
      },
      error: reject
    });
  });
}

private loadTransactions(): Promise<void> {
  return new Promise((resolve, reject) => {
    this.http.get<Transaction[]>('/api/transactions/recent?limit=8').subscribe({
      next: transactions => {
        this.transactions.set(transactions);
        resolve();
      },
      error: reject
    });
  });
}

Would I structure a growing Angular app this way forever? No.

For one screen, though, it is acceptable. There is no meaningful reuse yet, no routing, no write flow, no complex state model, and no product behavior that needs a separate service layer. Splitting it now would mostly create ceremony.

The important part is that this should remain a first-slice decision, not become the frontend architecture by inertia. When the application grows a second screen or a more involved interaction, the pressure to extract services and components will be real rather than theoretical.

What ran

The final application was manually verified end to end:

  • PostgreSQL starts using the provided configuration.
  • The Spring Boot backend starts successfully.
  • The Angular frontend starts successfully.
  • The frontend loads data through the REST API.
  • The API retrieves that data from PostgreSQL.
  • The UI displays the expected fictional account and transaction data.

The running Post 01 UI showing fictional accounts and recent transactions.

The backend also has tests for the first API behavior. The test artifacts from the reviewed state showed three Spring Boot test classes passing with zero failures or errors:

FinanceApplicationTests
AccountControllerTests
TransactionControllerTests

Those tests use H2 in PostgreSQL compatibility mode:

spring:
  datasource:
    url: jdbc:h2:mem:finance;MODE=PostgreSQL;DATABASE_TO_LOWER=TRUE;DEFAULT_NULL_ORDERING=HIGH

That is fine for this baseline, especially because the full app was manually verified against PostgreSQL. It is not the same as a PostgreSQL-backed integration test. If later work depends on database-specific behavior, that will need stronger coverage.

The defect I am keeping

Reviewing the result surfaced one confirmed defect, and I am keeping it in the baseline rather than smoothing it away.

The frontend package.json defines this script:

{
  "scripts": {
    "start": "ng serve --proxy-config proxy.conf.json",
    "build": "ng build",
    "test": "ng test"
  }
}

But angular.json configures build and serve; it does not configure a test target.

Running the command confirms the mismatch:

> personal-finance-frontend@0.0.1 test
> ng test

Cannot determine project or target for command.

This is a real issue. A developer sees npm test, reasonably assumes it is a supported verification command, and gets a broken command instead.

It is also a small issue. It does not affect the runtime slice. It does not invalidate the backend tests. It does not make the application architecture unsound. It should be recorded, not inflated.

I decided not to fix it before freezing the Post 01 baseline. A fix would be easy: either configure a real Angular test target and add the first frontend test, or remove the script until frontend tests are intentionally introduced. But applying that fix would change the evidence. The point of this experiment is to observe the first result, including the rough edges.

Engineering review

After reviewing the result, I would be comfortable carrying this first slice forward.

Not because the implementation is production-ready. It is not, and production readiness was not the objective of this experiment. For the requested first slice, the result is sound enough to keep:

  • It runs end to end.
  • It keeps product scope small.
  • It uses PostgreSQL for the actual local persistence path.
  • It creates an explicit Flyway schema and fictional seed data.
  • It exposes read-only account and recent transaction endpoints.
  • It renders the first Angular screen from backend data.
  • It includes backend tests for the seeded API behavior.
  • It avoids exposing JPA entities directly as API responses.
  • It avoids obvious overreach into excluded features or repository infrastructure.

There are contextual concerns to carry forward:

  • PostgreSQL-backed integration testing should be reconsidered when persistence behavior matters more.
  • Account retrieval is currently unbounded, which is fine for three sample accounts but not a permanent API strategy if accounts become user-created or large.
  • The Angular app should be decomposed when the frontend grows beyond this first screen.
  • The broken frontend npm test script needs to remain visible as a tooling gap.

Those are not reasons to reject the baseline. They are the next layer of engineering context.

What this shows, and what it does not

This experiment supports a narrow conclusion:

Given this repository, this stack direction, and this bounded prompt, Codex produced a runnable, reviewable first full-stack slice with one confirmed low-severity tooling defect.

That is useful evidence. It is not a universal result.

It does not prove that minimal guidance is always enough. It does not prove that AGENTS.md is unnecessary. It does not prove that repository instructions, scripts, CI, or specialized agentic infrastructure should be avoided. It also does not prove that Codex will stay disciplined as the application becomes more complex.

If anything, the useful baseline is more modest:

minimal repository + realistic bounded request
-> runnable first slice
-> reviewable engineering result
-> one small but real verification gap

That gives the series something concrete to compare against later.

The next question

Post 01 gives us a baseline implementation. The next useful experiment is not to add more features immediately.

This first request was clear enough to produce a usable slice, but it still left Codex to infer many things: project structure, test expectations, frontend verification, how much scaffolding to create, and what “minimal” should mean in practice.

That raises the next useful question: what changes when the request is shaped more like an engineering task instead of a broad feature ask?

The point is not to prove that more process is always better. The point is to find where extra specificity starts paying for itself.