Paul Solt Shares Workflow for Building Apps With Codex Sans Xcode
iOS developer Paul Solt describes an Agent Skill called AppCreator and a Makefile-based workflow using xcbeautify so Codex can build, test, and run iPhone and Mac apps with clean pass/fail output.
Original post · 6 min read
I have a new Agent Skill that will help you make apps. Without this skill you're going to waste a lot of time. Let me explain.
I was building a Dangerous Spider app with Codex when I noticed the agent kept missing its own build and test failures. It was looking at the wrong status codes. It genuinely couldn't find the error and would say everything worked (when it didn't).
Xcode compiler build output is extremely verbose. Actual errors get buried in thousands of lines of text. Trying to find an error in Xcode build output is like finding a needle in a haystack.
When you layer in agents, you're wasting time and context by asking them to find the errors.
Agents need to know what worked and what didn't work. It needs to be pass/fail, so I created an agent-designed workflow to do just that.
Here's the 7-step workflow I use every day:
1. Make Xcode Projects Agent Friendly with AppCreator
I built an Agent Skill called AppCreator. Run it once, and it scaffolds a new Xcode project or retrofits an existing one. Now your project is agent-ready.
At the heart of it: a `Makefile` wraps the CLI `xcodebuild` commands using xcbeautify. Clean, readable output from Xcode app builds and tests. The agent sees what failed, fixes it, and moves on. No verbose output to search. It works for iPhone and Mac apps.
Download and install the AppCreator Skill to make your app project agent-friendly.
2. make Is the Only Build Command I Use
The skill is designed so that one command builds and runs your app:
make
Your agent knows how to work with Makefiles, so this is just the starting point. You can extend it however you want.
I ask agents to set the default action to "build-and-run", and to use special build scripts so the freshly built app is always relaunched, just like Xcode.
In Codex CLI, type "make" to check the agent's work.
In the Codex app, set up a custom run action: Click on the Play button and set it to: make.
Finding the setting later requires a few more steps: Go to Environments → Project → View → Edit Local Actions → Actions → Set Action Script to `make`.
With a Makefile, you have a fast, repeatable way to build and run your apps (via the up-arrow on the CLI or the Run button in the Codex app).
3. Just Talk to It
If you want to get good results, you need to actively steer your agent.
Agents are good, but they'll do things you don't want, and the only way to prevent that is to steer the ship as they work.
I frequently double-check the work and redirect when I see agents doing the wrong thing.
I use Wispr Flow to talk out loud — describe what I want, how it should behave, what needs to change. After an agent hands back work, I start Wispr Flow as I play-test the new changes. This gives me the chance to talk through what works and what doesn't, and then, when I'm done, I can paste the transcript directly into the Codex app.
Plan mode is helpful for sparking thoughts about how features and edge cases. However, I have found that the plan mode isn't good enough.
Instead, I would use plan mode to help you think through the feature as a starting point. Use it to spark discussion so you can refine which features you actually build.
Software is nuanced, and if you take it feature by feature, you're going to get better software in the end.
4. Tests Keeps Agents Accountable
Without tests, agents can get sloppy. Tests allow agents to catch their own mistakes.
Ask Codex to write unit tests as it builds. Your goal is fast tests. UI tests are helpful for verification, but they can be annoying and slow to run. I like to have agents use UI tests to catch errors that are impossible to test with unit tests alone.
My recommendation is that you separate your tests into at least two targets:
make test
make ui-test
When an agent runs UI tests, it takes over your machine, which can be disruptive if you need to do anything else (This is where having a second Mac can be useful).
UI tests will slow everything down, so it's best to test them only at hand-off points. When your agent has wrapped up one task or several tasks. Don't do full regression testing for incremental work; instead, ask the agent to test only the smallest subset to verify their work and remain fast.
Have your agents read this article from Peter Steinberger, @steipete: Running UI Tests on iOS With Ludicrous Speed. Using that, you can help keep your test suite from becoming a bottleneck.
5. Log Runtime Results
Another indispensable tool is the use of logs and app artifacts. Have your agent add logging to your app. This gives it another tool for seeing what went wrong in real time.
Agents can stream development logs to a file or read them in real time to fix problems that are not immediately obvious.
When my agent repeatedly fails to complete a task properly, I know it's time to introduce more detailed logs.
Just ask your agent:
Please add verbose logs around XYZ so that you can see what is happening and fix the… continue on X ↗
