<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
<title>Joshua Valdez</title><id>https://joshuavaldez.com/</id><link href="https://joshuavaldez.com/"/><link rel="self" href="https://joshuavaldez.com/feed/"/><updated>2026-09-02T00:00:00.000Z</updated><entry><title>1.5 years with Claude Code</title><id>https://joshuavaldez.com/15-years-with-claude-code/</id><link href="https://joshuavaldez.com/15-years-with-claude-code/"/><updated>2026-09-02T00:00:00.000Z</updated><published>2026-09-02T00:00:00.000Z</published><content type="html">&lt;p&gt;It’s been a year and a half since Claude Code launched. In that time, coding with AI has gone mainstream. Anthropic and OpenAI have grown by leaps and bounds. And we’re seeing the complete transformation of a profession that’s only been around in machine form for 80 years.&lt;/p&gt;
&lt;p&gt;Last year I was addicted. I spent days and nights glued to my terminal building prototypes. I had already been AI-pilled and with Claude Code, I became an evangelist. .&lt;/p&gt;
&lt;p&gt;Fast forward to now and am I coding faster? Definitely. Do I enjoy it? I’m not sure. Do I still believe AI is the future of coding? Yes.&lt;/p&gt;
&lt;p&gt;Strangely, I feel more ambivalent about coding with AI now. The shift feels less radical as time has gone by. It’s no longer the future of coding, it’s just coding.&lt;/p&gt;
&lt;p&gt;Agents fill in the details of my work. I no longer have to bother with the intricacies of TypeScript.&lt;/p&gt;
&lt;p&gt;Yet, I still have to determine the direction of my work. Agents amplify direction. If you don’t know what you’re building, you can certainly get to nowhere fast.&lt;/p&gt;
&lt;p&gt;Early on when building with Claude Code, I compared it to being an engineering manager. I was happy to hand off the details. But I’ve now realized something is missing in that analogy.&lt;/p&gt;
&lt;p&gt;When I hand off a task to another engineer, I know that engineer will be responsible for it. I know they’ll consider their options, interact with others on the team and use their knowledge of our systems to make reasonable choices.&lt;/p&gt;
&lt;p&gt;I can prompt for the same thoroughness with AI and surely as the models get better, they’ll be more thorough.&lt;/p&gt;
&lt;p&gt;Sure we’ll have systems that handle reviewing a plan, making sure the code adheres to the plan, adheres to our codebase guidelines and accomplishes the goals we set out.&lt;/p&gt;
&lt;p&gt;But that sense that a task, a project, or an app is owned end to end is not there. Yet.&lt;/p&gt;
&lt;p&gt;I’ve seen ideas here referred to as a “software factory”. The software is churning out itself.&lt;/p&gt;
&lt;p&gt;For meaningful work that solves problems we care about, I’m skeptical.&lt;/p&gt;
&lt;p&gt;There’s something missing from AI-produced software that only a human can add. AI doesn’t produce out-of-distribution ideas.&lt;/p&gt;
&lt;p&gt;Humans will long be necessary to set direction and probe the limits of what’s possible.&lt;/p&gt;
&lt;p&gt;I don’t see that changing anytime soon.&lt;/p&gt;
</content></entry>
<entry><title>What changed?</title><id>https://joshuavaldez.com/what-changed/</id><link href="https://joshuavaldez.com/what-changed/"/><updated>2026-02-27T00:00:00.000Z</updated><published>2026-02-27T00:00:00.000Z</published><content type="html">&lt;p&gt;I’ve been using coding agents to write 100% of my code since Claude Code came out.&lt;/p&gt;
&lt;p&gt;The discourse lately about a big shift in agent ability initially surprised me. I agree the agents are great. Though I have not seen a huge difference between Opus 4.5 and 4.6 as well as Codex-5.2 and 5.3. The new models are incredible. The models just before were quite incredible too.&lt;/p&gt;
&lt;p&gt;I will admit, 4.6 and 5.3 seem to be just a bit smarter than their predecessors.&lt;/p&gt;
&lt;p&gt;So is the change in the discourse that these models crossed a threshold for people?&lt;/p&gt;
&lt;p&gt;Is the change that people had written coding agents off and came back to drastically better models?&lt;/p&gt;
&lt;p&gt;Or is it something else?&lt;/p&gt;
&lt;p&gt;I’d argue it’s a little bit of everything. The models are likely the majority, let’s say 70% of the change. But the harnesses also play a meaningful role.&lt;/p&gt;
&lt;p&gt;A few major improvements in the harnesses have made a huge quality of life difference in my day to day.&lt;/p&gt;
&lt;p&gt;Specifically, auto compaction has made a massive difference in being able to rely on an agent to complete it’s work. It can’t be understated how reliable auto compaction is a huge unlock to coding workflows.&lt;/p&gt;
&lt;p&gt;Claude Code’s auto compaction initially was borked and I never used it. I would dutifully create a new session once the context window filled up.&lt;/p&gt;
&lt;p&gt;Now, I have been working on greenfield projects for the past year. What about someone working on an old codebase?&lt;/p&gt;
&lt;p&gt;Autocompaction is in my view the single biggest improvement outside of the model. I played around with making a fork of Zed in December and something in the agent’s ability work work independently gave me confidence that I could actually do this. Auto compaction was a huge part of it.&lt;/p&gt;
&lt;p&gt;The ability of a model to take an underspecified prompt and do something reasonable with it has been a huge change. Early on with Claude Code, I would give it a task that was relatively well scoped and expect it to still need assistance. Now I can give it larger tasks, expect it to make reasonable decisions and test the output. That switch to larger tasks makes it more impressive to folks just tuning in.&lt;/p&gt;
&lt;p&gt;Finally, the models have been RLHF’d to ask questions and be more conversational. This tuning makes them feel fun to use! It’s like working with a real collaborator. Opus 4.6 shines particularly well in it’s conversational style. Codex-5.3 shines in it’s speed.&lt;/p&gt;
&lt;p&gt;I started to see the shift on conversational ability with Codex-5.1 and it it’s only gotten better since then.&lt;/p&gt;
&lt;p&gt;The tuning here to default to asking the user questions softens the hurdle for getting good results from the first prompt. In a sense, the models train the human users to disambiguate. That makes a world of difference in helping a human user understand how to get the best results from the model.&lt;/p&gt;
&lt;p&gt;A few other things I like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Claude Code’s use of color helps me grok ideas faster&lt;/li&gt;
&lt;li&gt;Codex’s app makes it seamless to start a bunch of ephemeral threads&lt;/li&gt;
&lt;/ul&gt;
</content></entry>
<entry><title>The quest for flow with AI coding</title><id>https://joshuavaldez.com/the-quest-for-flow-with-ai-coding/</id><link href="https://joshuavaldez.com/the-quest-for-flow-with-ai-coding/"/><updated>2026-01-09T00:00:00.000Z</updated><published>2026-01-09T00:00:00.000Z</published><content type="html">&lt;p&gt;I love programming with agents. I was an early adopter of Claude Code and I feel like it unleashed a burst of creativity for me.&lt;/p&gt;
&lt;p&gt;At the same time, I’ve been puzzled to find that as I use agents  more, I get this distinct sense that they’re slow.&lt;/p&gt;
&lt;p&gt;I’ve spent the past 6 months trying to root cause this feeling by building. At first I built my own cloud agents.&lt;/p&gt;
&lt;p&gt;My reasoning went: if agents are slow, I’ll write up a decent plan and then kick it to a cloud agent to finish while I move on to other things.&lt;/p&gt;
&lt;p&gt;Building the cloud agents was fun, in particular I enjoyed refining the interface and keyboard shortcuts for the cloud agents.&lt;/p&gt;
&lt;p&gt;At the same time, I found that the bottleneck then became twofold:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Testing the results of the cloud agents&lt;/li&gt;
&lt;li&gt;Iterating on a detailed enough plan for the cloud agents&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Testing the results is a hard problem. Codex and cto.new both attempt to solve this with loading a web app and screenshotting the results. That’s great for web apps but as any experienced software engineer knows, iterating on UI is only half the battle.&lt;/p&gt;
&lt;p&gt;So, setting aside that can of worms, I took inspiration from &lt;a href=&quot;https://repoprompt.com/&quot;&gt;Repo Prompt&lt;/a&gt; and built my own code/chat interface.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;/assets/images/the-quest-for-flow-with-ai-coding-cleanshot-2026-01-09-at-15-402x.webp&quot; alt=&quot;CleanShot 2026-01-09 at 15&quot;&gt;&lt;/p&gt;
&lt;p&gt;For a relatively small codebase, I could drop entire folders of code directly into the model’s context and ask for a plan to build the next feature I was working on.&lt;/p&gt;
&lt;p&gt;The plans I generated with this approach were great and I built an affordance to kick a finished plan into my cloud agents.&lt;/p&gt;
&lt;p&gt;So far so good.&lt;/p&gt;
&lt;p&gt;Except, not quite.&lt;/p&gt;
&lt;p&gt;The model often took a long time to formulate a plan. When you’re waiting 5 minutes for the plan, you can’t stay in flow. You no longer feel like you’re pairing with the model, you’re just waiting on it.&lt;/p&gt;
&lt;p&gt;Enter, GPT-5-Codex.&lt;/p&gt;
&lt;p&gt;Using this model felt like a revelation. It’s designed for coding and it strikes the right balance of being smart and fast to respond.&lt;/p&gt;
&lt;p&gt;I decided to transform my chat interface into a Codex driver and so far I’m very pleased with the results. It’s quite fun to use a chat UI to code.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;/assets/images/the-quest-for-flow-with-ai-coding-cleanshot-2026-01-09-at-15-082x.webp&quot; alt=&quot;CleanShot 2026-01-09 at 15&quot;&gt;&lt;/p&gt;
&lt;p&gt;At the same time, I haven’t quite answered my original question.&lt;/p&gt;
&lt;p&gt;How do I stay in flow while coding with AI?&lt;/p&gt;
&lt;p&gt;Do I lay out all my tasks up front and kick off simultaneous planning sessions? Do I spend more time writing elaborate RFCs for the agents to follow? Do I use &lt;a href=&quot;https://github.com/steveyegge/beads&quot;&gt;beads&lt;/a&gt;?&lt;/p&gt;
&lt;p&gt;I suspect we’ll be asking this question for a long time to come.&lt;/p&gt;
</content></entry>
<entry><title>Build your own background agent!</title><id>https://joshuavaldez.com/build-your-own-background-agent/</id><link href="https://joshuavaldez.com/build-your-own-background-agent/"/><updated>2026-01-08T00:00:00.000Z</updated><published>2026-01-08T00:00:00.000Z</published><content type="html">&lt;p&gt;&lt;img src=&quot;/assets/images/build-your-own-background-agent-screenshot-2025-08-13-at-11.webp&quot; alt=&quot;Screenshot 2025-08-13 at 11&quot;&gt;&lt;/p&gt;
&lt;p&gt;I recently shut down my startup, Engines. Coming off that experience, I felt a huge burst of creative energy to build and explore.&lt;/p&gt;
&lt;p&gt;I had some inklings that AI coding interfaces didn’t align with my sensibilities (mainly achieving flow while coding with agents) and I want to build my own background agent interface to see how it would work.&lt;/p&gt;
&lt;p&gt;Here’s my experience building a background agent on top of the Claude Code SDK with Claude Code itself.&lt;/p&gt;
&lt;p&gt;First I needed to choose a sandbox provider. Initially I started with Daytona but I struggled to get Claude running. It turns out that even though the Claude Code SDK runs headless, the underlying libraries used still require a terminal present. This posed an issue on Daytona but eventually I was able to figure out a workaround.&lt;/p&gt;
&lt;p&gt;Once I had chosen my sandbox provider, I chose a hosting provider. I’ve used Render in the past and liked it so it was the obvious choice.&lt;/p&gt;
&lt;p&gt;I decided to build the backend in Rust and the frontend in Typescript. Mainly because I’ve been enjoying using Rust for AI projects, since I find the compiler can catch many potential errors before they are committed.&lt;/p&gt;
&lt;p&gt;Building something end to end that worked took a week or so. Building something reliable has taken longer.&lt;/p&gt;
&lt;p&gt;The first roadblock I ran into was that Daytona did not support SSH from outside the sandbox. That was initially a critical requirement because I wanted to support SSH into the container so I could test things locally.&lt;/p&gt;
&lt;p&gt;So I switched to Modal. Surprisingly this was very straightforward since Claude had written a decent Rust interface for the sandboxes.&lt;/p&gt;
&lt;p&gt;Once I switched to Modal, I was enamored with how convenient it was and how clean its interface is. It’s incredibly fast. The startup times for its containers are incredible. Less than 10 ms due to two clever proprietary filesystems.&lt;/p&gt;
&lt;p&gt;I was able to quickly build a custom image for my background agents to have the right compile time tools.&lt;/p&gt;
&lt;p&gt;The most useful learning by far in this process has been to understand what background agents are useful for.&lt;/p&gt;
&lt;p&gt;I have come to see background agents as one of three things:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Exploratory agents for a new feature&lt;/li&gt;
&lt;li&gt;Quick bug fixers when I know the root cause or can point to a set of potential causes&lt;/li&gt;
&lt;li&gt;Executors for concrete plans&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I have come to the conclusion that for these tasks, the agent itself doesn’t matter much. I switch often between Claude Code and Codex and I find the results typically are about the same.&lt;/p&gt;
&lt;p&gt;That’s surprising! I would expect Opus to yield dramatically better results but I don’t find that to be the case.&lt;/p&gt;
&lt;p&gt;My take on the future is that background agents will become a commodity. Everyone will have one for the aforementioned tasks above and you will likely rely on a competent but not SOTA model to get things done.&lt;/p&gt;
&lt;p&gt;Since I was programming almost entirely with background agents, I came up with another tool to speed up planning. I found myself relying often on Repo Prompt (link) but I found the copy/paste tedious.&lt;/p&gt;
&lt;p&gt;I wanted to select files for the context window just like I would in Zed, with a Crtl + P workflow.&lt;/p&gt;
&lt;p&gt;So I built Swarm Desktop. It’s a chat interface that allows you to drop files into GPT 5 and chat with them. GPT 5’s context window is quite large so depending on your codebase, you can drop plenty of files.&lt;/p&gt;
&lt;p&gt;You can then ask for detailed plans based on your desired changes and get back a plan that Claude Code or Codex can follow.&lt;/p&gt;
&lt;p&gt;Swarm Desktop also allows you to spawn a task to Swarm Cloud with the click of a button.&lt;/p&gt;
&lt;p&gt;I find myself using this often when I discuss a plan at length with GPT 5 and then have it output the final plan for implementation.&lt;/p&gt;
&lt;p&gt;Through this all, I have found my preferred way to work with agents and I’ve built more of an intuition for when these tools are useful and when they need more direction.&lt;/p&gt;
</content></entry>
<entry><title>A startup is momentum</title><id>https://joshuavaldez.com/a-startup-is-momentum/</id><link href="https://joshuavaldez.com/a-startup-is-momentum/"/><updated>2025-11-13T00:00:00.000Z</updated><published>2025-11-13T00:00:00.000Z</published><content type="html">&lt;p&gt;Over the course of building Engines, I came to realize that a startup is momentum. You can build momentum in various ways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Outreach to customers&lt;/li&gt;
&lt;li&gt;Building new products and features&lt;/li&gt;
&lt;li&gt;Blogging about what you&apos;re building&lt;/li&gt;
&lt;li&gt;Raising funding&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To win, you have to do all these things. Not necessarily at the same time. But relentlessly over time.&lt;/p&gt;
&lt;p&gt;You have to tap into a deep well of demand for something people want and then harness that demand with relentless focus and execution.&lt;/p&gt;
</content></entry>
<entry><title>Stop reading, start doing</title><id>https://joshuavaldez.com/stop-reading-start-doing/</id><link href="https://joshuavaldez.com/stop-reading-start-doing/"/><updated>2025-09-12T00:00:00.000Z</updated><published>2025-09-12T00:00:00.000Z</published><content type="html">&lt;p&gt;If you try to read everything there is to know about a subject, you’ll never get started learning.&lt;/p&gt;
&lt;p&gt;Starting a company is an act of doing.&lt;/p&gt;
&lt;p&gt;When I first decided to become a founder, I did what I’d done when preparing for any other thing in life. I sat down to research how to become a founder.&lt;/p&gt;
&lt;p&gt;When brainstorming ideas, I would first search for competitors.&lt;/p&gt;
&lt;p&gt;When trying to sell to customers, I would first ask for feedback.&lt;/p&gt;
&lt;p&gt;But a startup is a process. It’s not a chain of thought.&lt;/p&gt;
&lt;p&gt;8 hours sending cold emails and 1 hour reflecting on the results is worth more than 9 hours “doing research”.&lt;/p&gt;
&lt;p&gt;Stop reading. Start doing.&lt;/p&gt;
</content></entry>
<entry><title>The unbearable slowness of AI coding</title><id>https://joshuavaldez.com/the-unbearable-slowness-of-ai-coding/</id><link href="https://joshuavaldez.com/the-unbearable-slowness-of-ai-coding/"/><updated>2025-08-19T00:00:00.000Z</updated><published>2025-08-19T00:00:00.000Z</published><content type="html">&lt;p&gt;I’ve been coding entirely with Claude Code for the past two months. At first it was exhilarating. I was speeding through tasks. I was committing like mad.&lt;/p&gt;
&lt;p&gt;Now, as I’ve built up a fairly substantial app, it’s slowed to a crawl. Ironically, the app I’m building lets me parallelize many instances of Claude Code at once.&lt;/p&gt;
&lt;p&gt;Often, I’ll have 5 instances running while I’m thinking about new features.&lt;/p&gt;
&lt;p&gt;The slowness comes in when I actually need to review all the PRs. One by one, I have to apply them locally. One by one, I have to step through the console logs. One by one, I have to tell Claude to fix the issues it created.&lt;/p&gt;
&lt;p&gt;Yes, it’s faster. I’m committing an incredible amount of code these days—more than I ever have.&lt;/p&gt;
&lt;p&gt;It also feels incredibly, maddeningly slow. Once you’ve felt that first boost of speed with Claude Code, you want every coding task to feel like that. &lt;a href=&quot;https://steipete.me/posts/just-one-more-prompt&quot;&gt;It’s addictive.&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Instead, as you build, you still have to serve as Claude’s QA engineer.&lt;/p&gt;
&lt;p&gt;Maybe one day we’ll solve that. But I’m skeptical it will be in the form of a CLAUDE.md. I can barely get Claude to consistently follow the bare set of rules I have, much less ensure it performs a complex integration test for a web app.&lt;/p&gt;
&lt;p&gt;Until then, I’ll keep pulling PRs locally, adding more git hooks to enforce code quality, and zooming through coding tasks—only to realize ChatGPT and Claude hallucinated library features and I now have to rip out Clerk and implement GitHub OAuth from scratch.&lt;/p&gt;
</content></entry>
<entry><title>On building trust</title><id>https://joshuavaldez.com/on-building-trust/</id><link href="https://joshuavaldez.com/on-building-trust/"/><updated>2025-08-05T00:00:00.000Z</updated><published>2025-08-05T00:00:00.000Z</published><content type="html">&lt;p&gt;I spent the last 30 days vibe coding &lt;a href=&quot;https://superswarm.dev&quot;&gt;Swarm&lt;/a&gt;. The experience keeps bringing me back to one overriding question about AI coding: how do I build trust in the agent’s outputs?&lt;/p&gt;
&lt;p&gt;Building trust in the outputs is critical to scaling usage.&lt;/p&gt;
&lt;p&gt;If I can check out a PR and be sure it’s correct by running a few tests locally, that’s great.&lt;/p&gt;
&lt;p&gt;If I can see the tests on the PR passed, that’s even better.&lt;/p&gt;
&lt;p&gt;If I can see the agent ran a comprehensive integration test in its own sandbox that ensures the end-to-end flow works correctly after the changes, that’s the best.&lt;/p&gt;
&lt;p&gt;I am a longtime Phabricator user, and one of my favorite aspects of the tool is the “Test Plan.” At &lt;a href=&quot;https://gem.com&quot;&gt;Gem&lt;/a&gt;, we used the Test Plan section to write down how we tested the changes in the PRs we put up for review.&lt;/p&gt;
&lt;p&gt;So far, I have not seen an AI coding tool that has a similar concept.&lt;/p&gt;
&lt;p&gt;I want to know the code generated is trustworthy. Trustworthy means more than just passing the linter, unit tests, and compiler. It means it won’t break production. It actually handles edge cases gracefully.&lt;/p&gt;
&lt;p&gt;With convincing test plans, we’ll be able to treat agents as actual engineers vs. the very smart interns they are today.&lt;/p&gt;
&lt;p&gt;Eventually, we will move beyond humans as the gatekeepers for trust and agents will be able to establish trust with each other.&lt;/p&gt;
</content></entry>
<entry><title>Engineering is writing</title><id>https://joshuavaldez.com/engineering-is-writing/</id><link href="https://joshuavaldez.com/engineering-is-writing/"/><updated>2025-08-05T00:00:00.000Z</updated><published>2025-08-05T00:00:00.000Z</published><content type="html">&lt;p&gt;I spent the last 30 days vibe coding and it’s changed how I think about programming.&lt;/p&gt;
&lt;p&gt;Programming is clarity of thought and writing. This was always the case but LLMs make it more clear.&lt;/p&gt;
&lt;p&gt;Once you’ve spent a lot of time exclusively using LLMs to code, they feel more like fancy autocomplete. They initially feel like a miracle. They can output so much code in so little time. Building a basic web app has never been faster.&lt;/p&gt;
&lt;p&gt;But they can’t replace clarity of thought.&lt;/p&gt;
&lt;p&gt;If you don’t ask for a specific structure in the code, they’ll put everything in one file. If you don’t specify edge cases, they’ll default to non sensical behavior.&lt;/p&gt;
&lt;p&gt;To get the most out of coding agents, you have to write very clearly.&lt;/p&gt;
&lt;p&gt;You have to sit down and really think deeply about what you’re building it before you start building it.&lt;/p&gt;
</content></entry>
<entry><title>Future of coding agents</title><id>https://joshuavaldez.com/future-of-coding-agents/</id><link href="https://joshuavaldez.com/future-of-coding-agents/"/><updated>2025-08-05T00:00:00.000Z</updated><published>2025-08-05T00:00:00.000Z</published><content type="html">&lt;p&gt;Soon we will all be using coding agents. Hundreds of them. The best software engineers will be the best multitaskers. The ones who can architect a great application, split it into parallelizable work and guide a team of agents to build it.&lt;/p&gt;
&lt;p&gt;Agents reward expertise and clarity of thought.&lt;/p&gt;
&lt;p&gt;We will still need software engineers. Our jobs will just be different.&lt;/p&gt;
&lt;p&gt;Companies will use a variety of coding agents just like companies use a variety of IDEs today.&lt;/p&gt;
&lt;p&gt;One team will prefer Claude Code. Another will use Devin exclusively. Another team will build their own agent, specialized to writing bullet proof Terraform.&lt;/p&gt;
&lt;p&gt;The coding experiences that win will deeply understand:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;How to build great UX so engineers understand the work agents are doing&lt;/li&gt;
&lt;li&gt;How to build great agents&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A winning experience must have both. Great agents without great UX will be increasingly harder to use as agents do more and more of the work.&lt;/p&gt;
</content></entry>
<entry><title>Principles for vibe coding</title><id>https://joshuavaldez.com/principles-for-vibe-coding/</id><link href="https://joshuavaldez.com/principles-for-vibe-coding/"/><updated>2025-08-05T00:00:00.000Z</updated><published>2025-08-05T00:00:00.000Z</published><content type="html">&lt;h3 id=&quot;1-you-get-out-of-it-what-you-put-into-it&quot; tabindex=&quot;-1&quot;&gt;1. You get out of it what you put into it&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;I’ve seen engineers say “AI doesn’t work well” and simultaneously be giving Claude Code prompts like “fix this.” If your task description to the agent wouldn’t be good enough for a human engineer, it’s not going to be good enough for an agent.&lt;/li&gt;
&lt;li&gt;Using agents is work! They can dramatically speed you up but you need to make use of them well.&lt;/li&gt;
&lt;li&gt;Building an intuition for what an agent will do well at and where it will struggle is critical.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;2-get-clear-on-what-you-want-to-build&quot; tabindex=&quot;-1&quot;&gt;2. Get clear on what you want to build&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Do you want to modify a React button to update on an app change? How do you want the button to change? What color should the button be? What state should the button store?&lt;/li&gt;
&lt;li&gt;If you can answer these questions, you will get a better response from the agent.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;3-use-o3-for-planning&quot; tabindex=&quot;-1&quot;&gt;3. Use o3 for planning&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;So far, o3 seems to be the best thinking model for coding. I dump my code into o3 in ChatGPT using Repo Prompt, ask for a plan that I can put in Claude Code, and then run the plan.&lt;/li&gt;
&lt;li&gt;If I am struggling with a bug, I ask o3 for information on the bug.&lt;/li&gt;
&lt;li&gt;Using Wisper Flow dramatically speeds up my prompting. I find myself giving a bunch of additional context to the model when I’m speaking instead of typing.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;4-test-agent-output-just-like-you-would-an-intern&quot; tabindex=&quot;-1&quot;&gt;4. Test agent output just like you would an intern!&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;You would not trust an intern’s code!&lt;/li&gt;
&lt;li&gt;You would pull the PR, test it thoroughly and review it.&lt;/li&gt;
&lt;li&gt;You should treat AI generated code similarly.&lt;/li&gt;
&lt;li&gt;You can speed up review with o3 before taking a look at the code to catch any obvious issues.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;5-you-are-the-architect-and-your-existing-codebase-is-the-guide&quot; tabindex=&quot;-1&quot;&gt;5. You are the architect and your existing codebase is the guide&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Do not trust the agent to architect well! You should ask o3 to plan good decomposition in advance and instruct it to structure things well.&lt;/li&gt;
&lt;li&gt;Your codebase being well decomposed will help here! If it’s clear to you where something should go, it will be clear to an agent.&lt;/li&gt;
&lt;li&gt;This is one of the biggest struggles for vibe coding. You can get started quickly but multiple Claude Code sessions later and your codebase is a mess. I often have o3 suggest a refactor which works quite well!&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;6-have-tests&quot; tabindex=&quot;-1&quot;&gt;6. Have tests!&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;This is one I am still struggling with. When I’m iterating quickly on a codebase, it seems superfluous to create integration tests.&lt;/li&gt;
&lt;li&gt;On the other hand, when an agent can run the tests, or at a minimum, run the type checker, it can usually iterate to a decent solution.&lt;/li&gt;
&lt;/ul&gt;
</content></entry>
<entry><title>1 week with Claude Code</title><id>https://joshuavaldez.com/1-week-with-claude-code/</id><link href="https://joshuavaldez.com/1-week-with-claude-code/"/><updated>2025-03-21T00:00:00.000Z</updated><published>2025-03-21T00:00:00.000Z</published><content type="html">&lt;p&gt;&lt;img src=&quot;/assets/images/1-week-with-claude-code-2025-engine-builder-repo-stats.webp&quot; alt=&quot;2025-engine-builder-repo-stats&quot;&gt;
&lt;em&gt;The repo stats for the Rust CLI agent codebase. 100% AI written LOC.&lt;/em&gt;&lt;/p&gt;
&lt;h3 id=&quot;hooked&quot; tabindex=&quot;-1&quot;&gt;Hooked!&lt;/h3&gt;
&lt;p&gt;Do you remember when you wrote your first working program? The high you got seeing it compile? The rush of changing a few lines and seeing it do something new?&lt;/p&gt;
&lt;p&gt;I first experienced that in middle school with my TI-89. I took the calculator manual and started to write simple programs.&lt;/p&gt;
&lt;p&gt;In high school, I wrote an entire stoichiometry program. I remember I was so addicted to coding that my chemistry teacher would ask me if I was following along with the lesson. Of course! I had already read the entire chemistry textbook cover to cover. So I told her what the lesson covered and got back to TI-BASIC.&lt;/p&gt;
&lt;p&gt;Now, hold on to the feeling. Because you’re about to feel it again. Claude Code is that good.&lt;/p&gt;
&lt;h3 id=&quot;claude-code-changes-everything&quot; tabindex=&quot;-1&quot;&gt;Claude Code Changes Everything&lt;/h3&gt;
&lt;p&gt;Claude Code has completely changed the way I write software. If you haven’t tried it, you should.&lt;/p&gt;
&lt;p&gt;Claude Code paired with Sonnet 3.7 is the best agentic software engineering experience I have tried.&lt;/p&gt;
&lt;p&gt;It can handle bug fixes, UI changes, test updates, and more.&lt;/p&gt;
&lt;p&gt;I’ve written an entire Rust CLI agent with Claude Code (and a bit of Devin). &lt;strong&gt;100% AI written. I have not touched a single line of code.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is my first time writing a lot of Rust, and I’m flying. I built the majority of the CLI agent over this past weekend.&lt;/p&gt;
&lt;p&gt;This is my workflow:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Spec out what I want to accomplish&lt;/li&gt;
&lt;li&gt;Write a detailed task message in Sublime&lt;/li&gt;
&lt;li&gt;Copy the message into Claude Code&lt;/li&gt;
&lt;li&gt;Have Claude iterate till the compiler passes&lt;/li&gt;
&lt;li&gt;Test it myself&lt;/li&gt;
&lt;li&gt;Report any errors back to Claude&lt;/li&gt;
&lt;li&gt;Repeat&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&quot;problems-with-speed&quot; tabindex=&quot;-1&quot;&gt;Problems with Speed&lt;/h3&gt;
&lt;p&gt;There is no way I can possibly review all the code that is generated by Claude Code. It’s too much. I’ve found myself getting carried away with the iteration loop and realize that I have a massive amount of unstaged changes in my git repo.&lt;/p&gt;
&lt;p&gt;That necessitated that I build &lt;code&gt;msg&lt;/code&gt; (also with Claude Code). It passes the diff of the git changes into Claude 3.7 and outputs a short commit message. This has saved me so much time and keeps me in flow.&lt;/p&gt;
&lt;p&gt;Additionally, the codebase is large enough that Claude Code can get a bit confused and break things between large changes. I’ve had to get more in the habit of:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Prompt Claude Code&lt;/li&gt;
&lt;li&gt;Test changes&lt;/li&gt;
&lt;li&gt;Commit&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If I miss the commit and the changes are directionally correct, Claude Code might break something already working.&lt;/p&gt;
&lt;p&gt;Another problem I have yet to address is how my cofounder will manage to review all the PRs I’ve been pushing. There’s just too much code! 😅&lt;/p&gt;
&lt;p&gt;It’s great for an MVP. It will be interesting to see how my workflow changes as we move this MVP into production use cases.&lt;/p&gt;
&lt;h3 id=&quot;the-vibes-are-real&quot; tabindex=&quot;-1&quot;&gt;The Vibes Are Real&lt;/h3&gt;
&lt;p&gt;Claude Code has forever altered how I relate to coding. It feels lighter. It feels faster. It feels like I’m floating.&lt;/p&gt;
&lt;p&gt;I can output an incredible amount of working code for $1 here or $1 there. Sure, Claude Code is more expensive than Cursor or Windsurf, but it’s also better. It produces great results at light speed.&lt;/p&gt;
&lt;p&gt;In the past week, I’ve spent &lt;strong&gt;$94.99&lt;/strong&gt; on Claude Code—&lt;strong&gt;$400 a month&lt;/strong&gt;—to output an incredible amount of software that I wouldn’t have been able to do otherwise. I sped through building a prototype and trying different ideas.&lt;/p&gt;
&lt;p&gt;What is the price of your time? Using Claude Code changed my mindset.&lt;/p&gt;
&lt;p&gt;Paired with Devin, I can achieve the equivalent output of a few engineers—worth hundreds of thousands of dollars. The only limitation is the speed of my own thoughts!&lt;/p&gt;
&lt;p&gt;This is not vibe coding. This is real, working software. I’m demoing it to a potential customer tomorrow; with a bit more polish, they’ll be able to use it.&lt;/p&gt;
&lt;p&gt;In a way, Claude Code helps me type code at a superhuman speed. So, instead of handcrafting each line, I am free to iterate on my startup&apos;s higher-order work: building what customers want.&lt;/p&gt;
&lt;h3 id=&quot;tips&quot; tabindex=&quot;-1&quot;&gt;Tips&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Make it clear to Claude how to run your tests&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Ideally, one simple CLI command: &lt;code&gt;cargo test&lt;/code&gt; or &lt;code&gt;npm run test&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Give Claude a linter!&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Be very specific with your prompts!&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Ask Claude to tell you its plan&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Spot issues with the plan that result from ambiguity in your instructions.&lt;/li&gt;
&lt;li&gt;Tell Claude Code to “think hard” or “think harder” to kick the Sonnet 3.7 model into thinking mode.&lt;/li&gt;
&lt;li&gt;This approach results in a more well-reasoned plan.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Ask for options&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Request various options and their tradeoffs when generating a plan.&lt;/li&gt;
&lt;li&gt;This allows you to choose the best option or instruct it on something it didn’t consider.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Have empathy for the LLM&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;LLMs are incredible, but they have non-human-like shortcomings.&lt;/li&gt;
&lt;li&gt;Develop an intuition for what will and won’t work through iterative testing.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Iterate on the outputs instead of the code&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Focus on the result you’re aiming for.&lt;/li&gt;
&lt;li&gt;Consider all the ways to verify the program’s results, encode these in tests, or have Claude Code write the tests.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;If you want to go all in, sign up for Devin too&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Typically, I kick Devin off right before I head to bed.&lt;/li&gt;
&lt;li&gt;I Slack it ideas before sleep and wake up to review and merge working PRs.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Also, consider giving &lt;a href=&quot;https://x.com/CodebuffAI&quot;&gt;Codebuff&lt;/a&gt; a try&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Similar to Claude Code and Aider, Codebuff has a streamlined terminal UI and doesn’t require you to approve every action, saving time on small fixes.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;im-not-going-back&quot; tabindex=&quot;-1&quot;&gt;I’m Not Going Back&lt;/h3&gt;
&lt;p&gt;After a week with Claude Code, I am never going back to writing hand-crafted code.&lt;/p&gt;
&lt;p&gt;There are a few areas where it struggles but I will iterate on my prompts. Specifically, it can’t build a great terminal chat UI. But that’s probably because what I want for the UI is underspecified.&lt;/p&gt;
&lt;p&gt;You have to overcome the initial inertia of being frustrated with bad outputs. And remember, prompting is a skill!&lt;/p&gt;
&lt;h3 id=&quot;engines&quot; tabindex=&quot;-1&quot;&gt;Engines&lt;/h3&gt;
&lt;p&gt;This post wouldn’t be complete without a word on &lt;a href=&quot;https://engines.dev&quot;&gt;Engines&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;We’re building a CI/CD agent. Drop it into any repository and it can:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Generate a working Dockerfile to containerize your codebase&lt;/li&gt;
&lt;li&gt;Run your linter automatically&lt;/li&gt;
&lt;li&gt;Run your tests&lt;/li&gt;
&lt;li&gt;Help you deploy your code&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If that sounds helpful to you, reach out to me at &lt;a href=&quot;mailto:josh@engines.dev&quot;&gt;josh@engines.dev&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id=&quot;additional-links-to-coding-with-llms&quot; tabindex=&quot;-1&quot;&gt;Additional links to coding with LLMs&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Simon Willison is one of my favorite newsletters to follow, and &lt;a href=&quot;https://simonwillison.net/2025/Mar/11/using-llms-for-code/&quot;&gt;his piece on coding with LLMs&lt;/a&gt; is worth a read.&lt;/li&gt;
&lt;li&gt;Ryan Endacott’s &lt;a href=&quot;https://ryan.endacott.me/2025/03/05/anyone-can-build-in-2025.html&quot;&gt;post&lt;/a&gt; has similar vibes and inspired me to write this.&lt;/li&gt;
&lt;li&gt;John Horton&apos;s &lt;a href=&quot;https://docs.google.com/presentation/d/1XRIflchTrZR2aqvxd7dROu8-76mf5N4Priwk8l9wNns/edit&quot;&gt;Claude Code presentation&lt;/a&gt; is great for tips! I learned about &lt;a href=&quot;https://git-scm.com/docs/git-worktree&quot;&gt;git worktrees&lt;/a&gt; (for multiple instances of Claude Code running at once on the same codebase)!&lt;/li&gt;
&lt;/ul&gt;
</content></entry>
<entry><title>On getting startup ideas</title><id>https://joshuavaldez.com/on-getting-startup-ideas/</id><link href="https://joshuavaldez.com/on-getting-startup-ideas/"/><updated>2025-03-04T00:00:00.000Z</updated><published>2025-03-04T00:00:00.000Z</published><content type="html">&lt;p&gt;Doing is better than thinking. Pick a direction, get a customer, iterate.&lt;/p&gt;
&lt;p&gt;We tried to take a top down approach to deciding what to build in multiple iterations of our company. Each time we approached starting that way, we ended up finding dead ends.&lt;/p&gt;
&lt;p&gt;It&apos;s impossible to tell why a company is truly successful from the outside. Their marketing is tailored to a specific problem they are solving for the customer so the marketing is downstream of the problem.&lt;/p&gt;
&lt;p&gt;To truly understand a problem, you have to talk to users.&lt;/p&gt;
&lt;p&gt;Talking to users means identifying the rough shape of a problem, asking people if they have it, and listening to what they say. They will guide you in solving their problem.&lt;/p&gt;
&lt;p&gt;This does not mean a vague &amp;quot;discovery&amp;quot; conversation. Pick a specific problem. Include a short description of the problem in your outreach. Then dive deep in the meeting.&lt;/p&gt;
&lt;p&gt;Once you know the shape of the problem, ask them to pay!&lt;/p&gt;
&lt;p&gt;If you never test willingness to pay, you will never know if your product is actually useful.&lt;/p&gt;
</content></entry>
<entry><title>Dev tools for agents</title><id>https://joshuavaldez.com/dev-tools-for-agents/</id><link href="https://joshuavaldez.com/dev-tools-for-agents/"/><updated>2025-02-20T00:00:00.000Z</updated><published>2025-02-20T00:00:00.000Z</published><content type="html">&lt;p&gt;A common theme I&apos;ve seen emerging lately is dev tools built specifically for agentic use cases. I think the next billion dollar companies in dev tools will be here.&lt;/p&gt;
&lt;p&gt;To clarify, I mean tools that LLM agents can use directly for the purposes of software engineering. Agentic tools are a broader category that includes web browsing, payments, etc. I think those will be big. However, I&apos;m primarily interested in tools that AI SWEs use.&lt;/p&gt;
&lt;p&gt;A few examples that resonate in this space are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://morph.so/&quot;&gt;Morph&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.runloop.ai/&quot;&gt;RunLoop&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://blog.replay.io/the-nut-api&quot;&gt;Nut API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.codegen.com/introduction/overview&quot;&gt;Codegen&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I think evals, memory and agentic frameworks will also play a role in AI SWEs. My take is that these will be AI SWE-specific, but I have yet to form a strong opinion on how these pieces will be tangibly different from the same infra for agents broadly.&lt;/p&gt;
&lt;p&gt;Weights and Biases is building an interesting trace-based approach to &lt;a href=&quot;https://medium.com/@shawnup/the-best-ai-programmer-from-weights-biases-04cf8127afd8&quot;&gt;AI SWE evals&lt;/a&gt;.&lt;/p&gt;
</content></entry>
<entry><title>Why I&apos;m starting a blog</title><id>https://joshuavaldez.com/why-im-starting-a-blog/</id><link href="https://joshuavaldez.com/why-im-starting-a-blog/"/><updated>2025-02-19T00:00:00.000Z</updated><published>2025-02-19T00:00:00.000Z</published><content type="html">&lt;p&gt;I come across a ton of launches and papers at the intersection of AI and software engineering. I&apos;m starting this blog as a way to keep track of what I&apos;ve read over time and also share my perspective on new developments.&lt;/p&gt;
&lt;p&gt;I&apos;m the founder of an AI SWE infra company: &lt;a href=&quot;https://engines.dev&quot;&gt;engines.dev&lt;/a&gt;. We&apos;re building the best platform for AI SWEs.&lt;/p&gt;
&lt;p&gt;I am a former infrastructure engineer from Uber and a former engineering manager at Gem.&lt;/p&gt;
</content></entry>
</feed>
