Rendered at 22:07:25 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
dbingham 5 hours ago [-]
I think most of this is correct, in spite of potentially being built on a bad assumption.
The assumption is that LLMs should be writing the code and human engineers reviewing and verifying the LLM output. And that this pushes the cost of producing down. And I fundamentally disagree with that.
Every time I ask LLMs to write code, even with Opus 4.8 (haven't tried it with Opus 5 yet), what I get ends up being totally rewritten. LLMs still aren't good at writing maintainable code. Can they write plausibly functional code? Yes. But it won't survive the long term. People using LLMs to write all their code are gambling on them eventually getting to a point where the LLMs can fix their own code. It's possible, but I wouldn't necessarily bet on it.
Where I have found immense value from LLMs is in code review. Repeated review by LLMs catches an amazing amount of potential issues. They really shine on security review, but are very effective with any kind of review.
The other thing that the "LLMs write code camp" misunderstands is that writing was never the bottleneck. Understanding was. And understanding the code is still the bottleneck. But understanding is truly gained during the writing loop. The understanding you gain from pure reading or code review is marginal compared to the understanding you gain while writing.
Most of the time previously spent writing was actually spent updating and deepening our understanding of the system under development. There's no replacement for that understanding in a world where LLMs are doing the writing.
But if you flip it: humans write, LLMs review, then you still get a major gain -- not in speed, but in quality. And you keep the understanding loop intact. I would propose that this might be the best way to deploy LLMs.
gste 2 hours ago [-]
I think this is outdated. If you follow spec-driven development, get the model to do all the planning work upfront, review and iterate the plan, write clear markdown file documentation on the abstractions and patterns you want to follow, then you have every opportunity to tell the model how you want it to write the code. If you use Opus or Fable 5 it will then write the code better and faster than you will.
tripleee 40 minutes ago [-]
you're a lot more skilled than I am if you're able to know the abstractions and patterns you want while only have a weak grasp on the actual code. For me the abstractions emerge as I understand what I'm actually working with
dalemhurley 14 minutes ago [-]
I think you can have some system design in mind upfront, but agree the bulk emerges as you code. Clean Code practices worked exactly this way in which you write and then rewrite the clean code.
preg_match 15 minutes ago [-]
I agree, but this works too because AI makes prototyping more rapid. It's common to be stuck with suboptimal solutions because of time constraints. Maybe the feature isn't really what the customer wants, or maybe the code isn't actually very good. But it reaches prod anyway.
But, with AI, the cost of code has gone down more than it already has. Well, cheap things are easy to throw away. So you prototype, prototype, prototype, and close the loop as much as possible with the customer. True agile development, not big A Agile.
The problem is this requires alignment from management, and we're just not seeing it at many company. They can't grasp that things have changed, and that throwing away code is free. They don't trust engineers to close that gap, so customers and stakeholders are still waaaaaay over there and we're delivering features they don't want.
eikenberry 2 hours ago [-]
IMO "get the model to do all the planning work upfront, review and iterate the plan" is backwards. It works better if you do an initial plan yourself and then have the AI review it (and iterate as necessary). Makes you think about the design for a bit so you can have a semblance of a mental model.
prymitive 1 hours ago [-]
I have never seen a plan written by LLM that wasn’t vague and light on details, every time I need to ask for more details and every time I ask to implement it trips over some dead end in the plan, discards the plan and continues as if there was no plan to begin with. Planning feels often just narrating the request and a wish list then what a human would do - methodically build enough understanding so that you’re confident of the direction you choose.
LLM plans are overhyped and overrated.
addaon 55 minutes ago [-]
This just doesn't match my experience writing idiomatic C code for embedded real-time systems with ChatGPT-5.5 and -5.6. Are Opus and Fable really that different? I general I find that a fully correct implementation for some function may take an hour and be nearly one-shot, but going from there to code I'd be willing to actually commit is at least 4x the time commitment with regular back-and-forth -- so four hours of deep engagement for every one hour of one-shot. ChatGPT-5.6 Sol seems to be better that -5.5 at the one hour do-it portion, but worse at the four hour do-it-well portion.
whatevaa 2 hours ago [-]
Sure. But try it on a project on non-trivial size and then tell us the cost of it. LLMS are billed per token.
bluefirebrand 16 minutes ago [-]
I haven't done "spec driven development" in my entire career, 15+ years
Until managers give up on Agile, we'll never get time to actually write specs
written-beyond 2 hours ago [-]
Depends on how long that planning/review cycle takes.
gste 14 minutes ago [-]
True, but I think people are also missing the trick behind markdowns that reference each other and are used for context window management. They go in your source control, they are reusable, extendable, and composable in the same way code is. So planning is not a one-off effort. It compounds over time. If you're doing it right your plan is referencing automation techniques - testing, CI/CD, etc. This further compounds so that the agents verify their own work against the standards that you set.
To be honest, the models are getting so good that they do most of this unprompted now.
theptip 2 hours ago [-]
Agreed, all of the staff/principal engineers I work with are ~99%+ AI generated code, and increasing business value delivered as a result. This is on planet scale infra not CRUD apps. (And yes, you do need to carefully review the output and give steers/corrections. It’s still faster.)
At this point if you can’t get the agent to write good code then either I) you are in a very specific niche (like Karpathy trying to write NanoGPT) that is extremely out-of-distribution, or II) skill issue, you need to learn how to prompt better.
It’s fine to have a skill gap! Just don’t delude yourself that the tools are bad and everyone claiming they are good is wrong.
piker 1 hours ago [-]
How is a GPT implementation “extremely out of distribution” but “planet scale infra” isn’t? That’s got me totally confused about your point that I was taking seriously.
potatolicious 1 hours ago [-]
A lot of stuff that underpins planet-scale infra is open source and has been for a long time. The general approaches and architectures for it have also been widely discussed and litigated in the commons so it's not just the code - the specific why is also very well documented.
So despite its importance much of it is actually pretty in-distribution.
piker 36 minutes ago [-]
Okay, but then that does not square at all with calling a GPT implementation out-of-distribution, let alone "extremely" so.
tripleee 35 minutes ago [-]
it's really cool how in the past measuring developer productivity was incredibly difficult, but now its become easy if someone's making a point about LLMs
oasisaimlessly 51 minutes ago [-]
btw, your last paragraph makes you sound like an ass.
lioeters 52 minutes ago [-]
> writing was never the bottleneck. Understanding was.
Well put. This insight is worth repeating in every discussion on the subject, from software engineering to mathematics.
The problem is that understanding is not the product being sold. The business model is for everyone to become consumers of what the magical genie generates, where the "understanding" is kept on the side of the model providers. This ensures a future generation of consumers dependent on someone else to provide the understanding.
Otherwise, you can create your own answers based on actually understanding the code, theorem, proofs, etc. Smart consumers of LLMs will use them to increase their own knowledge and understanding, so that the service is augmenting their intelligence, not replacing it.
SoftTalker 4 hours ago [-]
How do you handle the fact that an LLM can't seem to help itself from disgorging page after page of words no matter what it's asked to do? I've never seen an LLM say "this looks good as-is; I would not spend any more time on it; what's next?" it will always seem to suggest using another pattern or additional abstractions or other yak shaving.
But to be fair human code reviews have the same problem. It's like reviewers feel they have not done their job if they don't find something wrong.
YZF 2 hours ago [-]
I've totally had latest models just say "here are minor nits but this is good to ship" when reviewing code. Claude is a bit more verbose usually but Codex by default is pretty terse. My experience anyways. They also take pushback on suggestions (this isn't actually a problem because X) and they'll agree and say ship it (hopefully only when you're right and it's not a problem ;) ).
sheept 3 hours ago [-]
I personally rely on this behavior of LLMs. For example, for reviewing prose, I might ask the LLM to find five issues. If there are obvious issues I missed, then it'll find them. But if all the issues it lists are minor or hallucinated, then I can have some confidence there weren't any glaring issues.
I do the same for LLM code review comments: some changes are out of scope or could be moved to a separate PR; some edge cases don't happen in practice and should just fail noisily instead of writing more code to maintain. When these are the only issues it's raising, then I know it's done.
RyanHamilton 4 hours ago [-]
I've tried and tried to get an LLM to delete code, it couldn't do it. I knew a file was 40% bad so what I ended up doing was deleting the file, then asking AI to create the missing file. That's how I got AI to delete code :)
bluefirebrand 14 minutes ago [-]
Sounds like the AI got you to delete the code for it.
clbrmbr 3 hours ago [-]
It is surprisingly hard to do in a single prompt. I’ve had good luck though with being very explicit about asking for review, then to think deeply about what actually matters, and then reduce to a very short extremely concise reply. —- the model NEEDS to output those pages of text as part of its thinking process. It’s used to doing it in the assistant reply but can be cajoled into doing it in the thinking tokens.
dbingham 4 hours ago [-]
Yeah, that's just a problem with review. I tend to keep it going until the stuff its finding is stuff that I look at and think "I can live with that". Usually by that point it's found the worthwhile stuff and what remains is minor or, like you said, yak shaving.
SoftTalker 4 hours ago [-]
There used to be a bit of "wisdom" passed around that said "when you submit something to the client (or management, etc) for approval, leave an obvious mistake somewhere." Then they will find it and point it out, you can easily correct it, and everyone is happy. Otherwise they will find something to critique, just to demonstrate that they contributed something to the review.
I wonder if the same trick works with AI?
clbrmbr 3 hours ago [-]
I think so, if you can inject synthetic errors with the same distribution as the errors you care about.
mike_hearn 2 hours ago [-]
What tools are you using? I've found that Codex will happily tell me there's nothing more to do after a few rounds of /review and fix (or sometimes, it picks up on low value points that are technically correct but not worth changing).
wuschel 3 hours ago [-]
My guess would be that this is a matter of fine-tuning the model to the problem. You would need to run the review experiment N times, observe the distribution of answers, and then fine tune with model operational parameters.
Perhaps someone more knowledgable could jump in here to clarify?
chucky_z 4 hours ago [-]
I’ve found stating “only call out critical issues, report but ignore everything else” will basically get what you want. Put that in a review skill and you’re good to go.
To be clear I mostly only use Opus and Gemini Flash but this might work for others too.
chrisjj 3 hours ago [-]
> I’ve found stating “only call out critical issues, report but ignore everything else” will basically get what you want.
Disabling your C compiler warnings works too. You get to ship then leave work early!
wonnage 3 hours ago [-]
The verbosity itself seems to be a problem with post-4.6 Claude rather than a general issue with all LLMs. and IME the yak shaving is due to an overly generic prompt. We have a generic automated review bot but it’s prompted to only look for errors and never suggests refactoring; that was a conscious trade off to avoid what you’re talking about. If you’re reviewing interactively then you can give additional guidance about any code smells etc. that jump out
clbrmbr 3 hours ago [-]
I’ve found that doing the full requirements capture, planning, writing, reviewing, gardening loop with frontier models has worked quite well since last October, and phenomenally since Fable 5.
The key I’ve found is human peer review. The reviewer jumps on a live call with the developer, pulls up the PR with transcription on, and asks questions. At the end of the call, the transcript passes back into the coding agent and the PR is polished up, becoming more self-documenting, and the humans are left with some degree of common understanding of what’s going on.
I’ve been operating my team of ~15 this way for 9mo to great effect… there is simply no going back to the stone ages.
oogali 3 hours ago [-]
Here's an odd request for you, as a fellow New Jersey-an...
Can I watch/observe one of your review sessions?
Every few weeks, I hear the beginnings of a great approach towards working with LLMs but I rarely see it in practice.
If you're open to this, remote or in person, ping my username at gmail.
luaKmua 56 minutes ago [-]
This is where I've landed too, I imagine based off all the responses that it depends on domain and also the rigor you care about in your implementation. For me, what I do, and the quality required, there's just not much time (if any) saved on handing off the implementation to an LLM.
But that doesn't mean they're useless either. I use them all the time for review as you mentioned or to knock out one-off scripts that don't go anywhere near source control. There's just no world where I don't need to understand every line of code that I'm responsible for getting into our project.
tracerbulletx 3 hours ago [-]
I have not experienced this "llms dont write maintainable code narrative" I just tell it the shape of the entities and apis I want vaguely and the mental model and the output is excellent.
supriyo-biswas 2 hours ago [-]
I could do this in the beginning of 2025. Unfortunately, given the over-proactiveness of model these days, I find that they end up inferring what my original request was and implementing it anyway, and then writing unit tests, etc. even when all I asked for was to wire up a few components, as an example.
lukeschlather 1 hours ago [-]
I actually find the converse, that I try to scope things down into chunks small enough for the model to work on, but it will start trying to hack off pieces that don't belong because it thinks the restricted scope is the whole world and I have to stop it and explain the bigger picture so it stops trying to remove things that serve the full scope.
ManuelKiessling 25 minutes ago [-]
To be honest, it feels like you are contradicting yourself.
If LLMs are, as you stated, really good at catching potential issues, then they are, almost by definition, really good at producing code without potential issues, if guided correctly: all they need to do is inspect and iterate, until they do not find any more potential issue in the code they produced.
lordnacho 3 hours ago [-]
> The other thing that the "LLMs write code camp" misunderstands is that writing was never the bottleneck. Understanding was. And understanding the code is still the bottleneck. But understanding is truly gained during the writing loop. The understanding you gain from pure reading or code review is marginal compared to the understanding you gain while writing.
This was my stance a couple of years ago, but now I've given it up.
It turns out writing actually was the bottleneck. You can understand perfectly well what you want, but writing it is long and tedious to the point where you find excuses not to do it. Particularly with version 2, the step where you have an OK system and you want to improve it. Quite a lot of changing the code is just useless busywork: re-wiring old functions, moving imports around, searching for locations that benefit from extracting a common piece of code. And each time you do one of those, there's a decent chance you did something even more trivial like forgetting a semicolon or calling the wrong function.
Now that I have an LLM helping me, I can see why. The critical decision is a terse declarative like "we need to have several TCP connections instead of one, and just use the sequence number to arbitrate". A human junior programmer could perfectly well understand what this meant, but he would have to go through all of the above to get to the final product. Now, I can just tell the LLM and I will get what I want, even with the things I didn't explicitly state, without spending attention.
This means I can use my attention on the things that matter. So instead of spending today thinking about how to arbitrate between the TCP connections and tomorrow thinking about pre-calculating my outgoing orders, I can just do both today. I don't waste the good waking hours chasing minor bugs, I just think about the large structure.
I get the feeling the best programmers of years past were actually masters of the little things, which led them to be able to look at the big things. Essentially it was cheaper for them to get to the top of the mountain, where you can see the landscape. Kinda like how the kid who was good at mental arithmetic in primary school was also good at calculus at the end of high school: if you don't have to concentrate on the little things, you have time for the big things.
twister2920 3 hours ago [-]
> And each time you do one of those, there's a decent chance you did something even more trivial like forgetting a semicolon or calling the wrong function.
how did you decide to pick the most trivial kind regression for this example? do you compile your code before checking it in?
> A human junior programmer could perfectly well understand what this meant, but he would have to go through all of the above to get to the final product. Now, I can just tell the LLM and I will get what I want, even with the things I didn't explicitly state, without spending attention.
the main efficiency you have described here is offloading the verification of a change onto the LLM. that is the bottleneck. readers can decide whether a non-deterministic statistical model is a good tool for this job
> I get the feeling the best programmers of years past were actually masters of the little things, which led them to be able to look at the big things
the best programmers understand that their job is to automate workflows, and that includes their own. if you're worried about missing a semicolon, I'm sorry to say that's a skill issue
lordnacho 2 hours ago [-]
> how did you decide to pick the most trivial kind regression for this example?
Why would this be a regression? You might just be writing a new line of code.
> do you compile your code before checking it in?
Well obviously. That is generally how you discover that a semicolon is missing.
> the main efficiency you have described here is offloading the verification of a change onto the LLM. that is the bottleneck. readers can decide whether a non-deterministic statistical model is a good tool for this job
No, it's the time between you deciding something needs to be done, and it being done, that is the bottleneck. You cannot avoid trying to compile the code and testing it. Now you can get to that test without paying attention, which is time you can use productively.
> readers can decide whether a non-deterministic statistical model is a good tool for this job
Somehow, the non-deterministic model has built me the deterministic code that I want, very fast, pretty much all the time. A year ago it would get stuck. Now it doesn't, for me at least, and for competent programmers that I know.
> the best programmers understand that their job is to automate workflows, and that includes their own. if you're worried about missing a semicolon, I'm sorry to say that's a skill issue
Well yeah, and I've automated my workflows completely. I don't have the problems I used to have. If you haven't caught on to the new way of working, well, that's a skill issue...
maccard 1 hours ago [-]
> Somehow, the non-deterministic model has built me the deterministic code that I want, very fast, pretty much all the time. A year ago it would get stuck. Now it doesn't, for me at least, and for competent programmers that I know.
I still find the models get stuck or go on _massive_ side quests. Just today, I asked claude to write a hello world C++ program using import std; I interrupted it when It decided I needed a new toolchain installed, and started checking for docker installations. This is super basic stuff, it hadn't even generated a plan, it just started searching for LLVM versions rather than running clang --version.
> If you haven't caught on to the new way of working, well, that's a skill issue...
Honestly, it feels like the emperor has no clothes on this topic, and the crowd defending LLMs to death are way too quick to call it a skill issue.
lordnacho 7 minutes ago [-]
> Honestly, it feels like the emperor has no clothes on this topic, and the crowd defending LLMs to death are way too quick to call it a skill issue.
I feel it's the other way around. The LLM skeptics are unwilling to admit that these things can get you there faster than you would on your own, in the face of clear evidence.
twister2920 1 hours ago [-]
> Well yeah, and I've automated my workflows completely. I don't have the problems I used to have. If you haven't caught on to the new way of working, well, that's a skill issue...
I use LLMs, but they're just a tool in the workflow, and I make sure to review the output. they might remember semicolons but they make much more pernicious mistakes that are harder to detect
lordnacho 9 minutes ago [-]
But are those pernicious issues more common than what you would write yourself? Keeping in mind that you will get your LLM code back a lot sooner and have more time left over to check for them?
Valectar 2 hours ago [-]
> Quite a lot of changing the code is just useless busywork: re-wiring old functions, moving imports around, searching for locations that benefit from extracting a common piece of code.
If other people are dissatisfied with LLM output quality while it seems to work fine for you, you might want to consider that the quality of code you produce is closer to the quality of code the LLM produces than what those other people are producing.
What you posted there, for example, about most of changing code being busy work is a pretty big red flag for a codebase. One of those "large structure" things that you're supposed to be paying attention to is the architecture of the code. There's always the chance that some change you need to do goes against the grain of the solution you architected, and you need to make changes all across your codebase to fit it in, but in general the point of modularity and good architecture is that when you make a change you just have to make that one change, ideally just changing the logic of the one responsible function with only minor changes required anywhere else in the codebase. If you're consistently having to hunt throughout the code for related functions that you need to rewire that's a sign that your architecture does not fit with the direction your codebase is evolving, or alternatively that you don't have much of an architecture to begin with and your code is highly interconnected.
Actually one habit you mention at the end of that quote can worsen this issue: "searching for locations that benefit from extracting a common piece of code". Tautologically this is a good thing as you define it as only working on locations that will benefit, but given the frequent need for rewiring of functions I would hazard to guess that you've "deduplicated" code a bit overzealously. Just because two functions share some common code does not necessarily mean it is appropriate to pull that out into a function. Deduplicating is good if conceptually the code is a single thing that you would always want to keep in sync, as it means that when you need to make a change to it you don't have to hunt down all the places it's used. On the other hand, if you find yourself frequently needing to delve in to these functions to rework them because you need to make a change to how it's used by just one caller, your "deduplication" has added to your workload, and probably created some overcomplicated code in the function that is in reality handling multiple distinct needs.
I hope this doesn't come across as too condescending, and if I've just wasted your time explaining principles you already understand I apologize. I don't know you or the code you're working on so I can't exactly confidently judge your work solely on a few paragraphs. It's just that your mention of how your experience of coding has been different from what others have described, and specifically that, for you, writing has been the bottleneck rather than understanding, combined with the specific issues you describe facing, imply to me that you may not realize that the approach you are taking to producing code yourself may be significantly different from how other Software Engineers are producing code, and that may account for some of the differences you note in your personal experiences programming.
lordnacho 1 hours ago [-]
I understand what you mean. Those are just some small examples. But I had this conversation with a friend today, and it goes something like this:
1) It was good for me to spend years learning the little stuff. Loops, variables, if conditions, how to import stuff, git, debugging things, reasoning about the flow of control. Classic coding.
2) I had a false dawn at about 10 years in. I thought I understood a lot.
3) I learned I had a lot to learn. Very wide areas of programming I'd never touched, ways of thinking that started to click.
4) I spent another ten years covering holes, building a different type of experience. My guesses about how to do a project are much better now. My guesses about what really matters have changed.
5) Now the small stuff is actually just bothering me. I'm not going to learn much more from staring at little things. There are larger architectural things to think about, and the little things are just friction.
So that's where I'm coming from. I get that a lot of pushback is going to be from 10-year-me, who thought he'd gotten to a high level of understanding by slogging through the little stuff.
leptons 3 hours ago [-]
>It turns out writing actually was the bottleneck. You can understand perfectly well what you want, but writing it is long and tedious to the point where you find excuses not to do it.
Speak for yourself. Writing code has never been a bottleneck for some of us. I can't speak for everyone, and neither should you.
>Now, I can just tell the LLM and I will get what I want, even with the things I didn't explicitly state, without spending attention.
This should worry you. All too often the LLM invents things I didn't ask for and implements things I didn't need. YMMV, I guess. If slop gets the job done, and nobody notices, then who should care?
dwayneII 5 hours ago [-]
I write code in Go in an established codebase. 90% of the time the LLM gets a basic api/feature right and fully e2e tested as long as I give it enough business context in the prompt. The other 10% of the time I have to do some follow up prompts to either change the behaviour or change the approach of a given step in the flow or just to point out that the wrong pattern was used and please rather use our codebase standard. I haven’t actually hand-written code in well over a year. Is this way faster? Absolutely. Does it lead to better code? Yes, because I have time to write all the regression tests that keep the behaviour as I wanted it when others come bungling around.
sanderjd 5 hours ago [-]
I do think the comment is right about understanding. I'm still feeling out what I think is the right level of understanding to invest in now. I think it is probably not the deep line by line understanding I used to have of the systems I worked on. But I also think it's easy to remain too aloof from how the system is being built, and that this is very bad. I think the right answer is somewhere in between, but I'm still working through my own process and forcing functions to strike that balance properly.
YZF 2 hours ago [-]
In my domain since circa Opus 4.5 LLMs write code as good as most engineers given well defined small enough chunk of work. They refactor. They write tests. These days LLM also debug/troubleshoot better than most engineers.
Are they as good as handcrafted code by 0.1% of top software engineers. Generally no. But neither is 99.9% of real code.
LLMs also are good at code reviews. What they'll miss is often the big picture but they can still catch plenty of issues. I still want to see a human in the loop in my domain.
Totally agree that writing the code was never the bottleneck. We're not seeing massive productivity gains even if some code is written faster. It's not just about understanding but also various other activities that happen in large companies and teams.
Also agree LLMs can be used to gain quality but realistically most orgs are going to aim for "fixed or decreasing" quality at lower costs.
himata4113 3 hours ago [-]
There is the flip side of where the code is no longer read by humans and is becoming the prevailing way software is shipped in tiny businesses. You don't need to code to be maintainable since you will never maintain it, the AI will and the quality will naturally improve as models improve. For example 5.6 sol and fable are showing signs where you can feed garbage in and it will spit out something pretty decent, definitely not the quality you'd expect from a senior developer with millenia of experience, but that of your average grunt worker turning words in an issue board into code.
However, sometimes then I tell it to write an app with detailed instructions and it spits out garbage so your mileage might vary.
lukan 5 hours ago [-]
" But understanding is truly gained during the writing loop. The understanding you gain from pure reading or code review is marginal compared to the understanding you gain while writing."
Debugging code step by step is how I understand complicated code.
drTobiasFunke 2 hours ago [-]
“Not in speed, but in quality” — unfortunately though, not a single CEO, executive, company, VC, investor or anyone with the power to make decisions cares about quality instead of speed.
consumer451 2 hours ago [-]
I agree. Also, what does "code quality" even mean in the age of agentic dev? If the code works, and is secure, what else matters?
I do know what good code looks like, but does that even matter anymore? All I know is that now, I get to focus on endless UX polish, which is the only thing the matters.
I feel like we are living through something like the Protestant Reformation, where priests once spoke Latin, and then started to speak in plain local language. The old guard did not like this.
ilovefood 4 hours ago [-]
You're right about understanding. Where we might disagree is the conclusion. I think it's possible to build the understanding without having to type out syntax by hand (debugging, writing tests, etc). Maybe we're missing much better verification tools? LLMs will likely play a big part in those too (explain the codebase, walk me through A, B, C etc)
bcrosby95 4 hours ago [-]
I've come to realize that the smarter LLMs get, the less they understand the point of abstractions.
I've been using it for a unity game for the past few years. Nowadays it will go sleuthing into packages and assembly and make decisions based upon what it sees there.
It will make comments about why it's doing something based upon a function call 3 methods deep.
God forbid any of these details change in a minor version update.
hkpack 3 hours ago [-]
Try reducing thinking and use lighter models. For example I observed that using Sonnet works much better (compared to Opus) for tasks I want to be in control of architecture and just need a faster code input.
broast 5 hours ago [-]
many of us work with swarms of cheap off-shore contractors and have to rewrite or finish their code already as it is. The llm is much lower friction
rufius 4 hours ago [-]
IME, the LLM when used well also gets closer to correct.
vatsachak 2 hours ago [-]
I hate saying this but unironically opus 5/gpt 5.6 are writing maintainable code on the order of 500-1000 lines
win311fwg 3 hours ago [-]
> LLMs still aren't good at writing maintainable code.
I find that depends on the target language. They can be good at writing maintainable code, but not consistently across every language.
The languages beginners usually gravitate towards are especially hard for LLMs to produce quality output for. Presumably this is due to the training data including all the unmaintainable codebases written by beginners in those languages, which hasn't allowed the LLM to converge on recognizing what a maintainable codebase looks like in those languages.
marginalia_nu 5 hours ago [-]
Code was very rarely the bottleneck in the first place.
If programmer productivity was something we actively optimized for, we wouldn't have crammed programmers like sardines in warm and noisy open floor offices with 2000 ppm CO2 levels and then further constantly interrupt them with emails and slack pings and meetings all day long, Jira rigmarole wouldn't make up a significant portion of what they did, programmers would have instead mostly been thinking and programming.
We've always had the ability to 2X if not 10X the output of each and every one of those poor souls. You don't end with this sort of programming purgatory because it's a productivity optimum, it very clearly isn't, but because it's a billable hours optimum and/or an org chart clout optimum and/or because of Jevons paradox got hands even in business management and the IT department was allocated too many dollars.
raverbashing 5 hours ago [-]
> Code was very rarely the bottleneck in the first place.
I disagree. Kinda
What AI has made much simpler is that you don't have to waste time checking docs and have the best autocomplete system by a long shot - this was a bottleneck unless you were doing Java or some other language with "perfect" AC
What AI made "kinda easier": solving for usual problems. The stuff you would search Stack Overflow, or think a couple of minutes for an optimized solution - not a bottleneck but not 100% smooth neither
You still have to test and validate your code. AI made this easier-ish but this is still where I see manual work being needed (even if you are automating tests - you still have to think on what you want the code to do)
sdevonoes 26 minutes ago [-]
But searching on google or SO wasn’t the bottleneck either. It was searching on your own codebase (and dependant ones) for the exact place (a class, a file, a function) where to put/delete/update business logic without breaking production.
LLMs help, but they haven’t been trained on our own repos. I don’t need the LLM to help me with algos that are available online… I need them to help me with custom business logic
bitlad 5 hours ago [-]
I disagree. Coordinating 100 folks to align on a problems statement with "agile" and sprints was always the biggest bottleneck. Difference in opinion, internal politics is always a bottleneck.
I think now, code is the bottleneck. Just because you can generate million lines of code, people with different skill level think they are accomplishing the task, testing, merge conflicts, trust has become the bottleneck.
parpfish 4 hours ago [-]
A pattern I have seen:
The “old way” would be lots of debate (both bike shedding and useful) among engineers during design phase, and then you’d implement.
Now it’s shifted so there are no design docs and there is only the generated prototype. People trying to do their design review while there’s already a functional-ish prototype and it goes nowhere. There’s an anchoring effect in place because the first thing already exists and management says “this seems to work, just use it and move on”. The result is that useful debates about substantive issues don’t happen and bikeshedding is all way get to do
drTobiasFunke 2 hours ago [-]
This is so accurate. I am a fairly senior individual contributor and its been more than a year since I saw a good design doc or quality design discussion. Before you can even question the design, someone has generated a 50k line prototype and already made up their mind because of all the “you’re absolutely right..” and “X is exactly what you need..” from AI.
skydhash 53 minutes ago [-]
This very much. First thing built is what get approved. Even if it will generate 10x the workload down the line, because today's quality bar is very much demo level.
And then you get paged at 2am because prod is down and the support channel is more active than the team's one.
fragmede 4 hours ago [-]
The distance between substantive issues vs bikeshed sized issues is non-linear and ill-defined. Sometimes: yeah, we're just bikeshedding. Other times, it's important to suss out unshared implied context that doesn't match between various stakeholders and will have outsized ramifications later. What's new is the speed of executing changes. If there was no substantive debate about what programming language to even use so everyone is equally happy (read: sad), so what? Have the LLM rewrite the entire project in another language over the span of a couple of days after that discussion is had because the language used has specific shortcomings that have been unmined.
marginalia_nu 5 hours ago [-]
This just says that we can output code faster. I'm saying that the rate at which we output code wasn't the thing that was slowing down development. We've always been able to increase that even without AI by adjusting the working environment and removing obstacles to programming.
In larger organizations, quite often it's the business that is holding back development. They can only handle so much change and speed needs direction to be velocity. Drafting requirements is generally much slower than implementing them.
Like the number one complaint from programmers has been that they don't get to do programming. They want to write code, not update jiras or spend hours in meetings.
baron3dl 5 hours ago [-]
What is the right, best software organization in the current era of AI coding? This question is critical and wholly unanswered in comprehensive research along the same axis as Accelerate (2018, Forsgren, Humble, Kim).
There are a lot of (excruciatingly) long-form posts about what folks are pioneering but not a whole lot of follow up about what failed. Where are the short posts on the negative space? How did halving your staff work out? Flattening your org? All those dark factories, what haven't they produced? How about all the other things tried, failed, and unceremoniously scrapped?
We need to explore and communicate the negative space more efficiently. Don't repeat the same mistakes, and don't make me read 2653 words when 300 do it better.
VeninVidiaVicii 5 hours ago [-]
Hard to say what caused what, but the internet seems to mistake verbosity for authority, and so does AI.
baron3dl 4 hours ago [-]
A tech comm course I took in college was graded on two 20-page papers and accompanying 5-minute presentations. That was like 10k written words in a single semester. It was a challenging class, and gave me substantial sense of accomplishment, just to hand in completed work.
Similar to a functioning side project in the 5-10k LOC range. Announcing something that worked a year ago, was laudable, even if not profitable.
I vibe coded 15k LOC this morning and read 20k words of AI generated text while doing so. No longer are either noteworthy or valuable public contributions just by virtue of having been done. I don't think that's widely recognized yet.
w10-1 2 hours ago [-]
> mistake verbosity for authority
Isn't that the essence of "attention is all you need"?
sanderjd 5 hours ago [-]
Great framing.
coffeebeqn 4 hours ago [-]
> and don't make me read 2653 words when 300 do it better.
A million times this. Can we please RL the next models to learn the “if I had more time I would’ve written a shorter letter” method please.
I see it every day in tickets, many communication channels, PR descriptions, comments, documentation. All have at least 70% verbose fluff which is so taxing and makes it very hard to keep track of the one important thing they’re trying to communicate in the message
mgaunard 5 hours ago [-]
The cost of code actually increased; code debt is being accumulated faster than we can clean it up.
VeninVidiaVicii 5 hours ago [-]
Yeah but look how much there is! Aren’t you impressed?
leptons 3 hours ago [-]
We still pay our software developers the same amount, but now an AI is the middleman taking more and more money (in the form of tokens) every day. With every new model, more tokens are consumed so the new model can "think" more. Nobody's even really reading or reviewing the code the AI produces, so software quality goes down, and we're paying more to produce it. And when there's a serious problem with the code? No human is going to dig through the mess the AI created - so burn even more tokens/money trying to get the AI to fix it, or just start over from scratch again with the AI, hoping for better results. It's the definition of insanity. I'm looking for a new job, maybe even a new industry, a new career path - after 30 years in this industry, software development is jumping the shark.
thih9 39 minutes ago [-]
We’ve already lived through multiple floods of unmaintainable code, eg with unstructured JQuery.
New tech stacks will appear and they will handle the mess to some extent.
chickensong 2 hours ago [-]
> Sorting requires a model of how your org actually behaves: trust relationships, hallway knowledge, the consequences of past decisions. Almost none of this is written down.
This is a long-standing problem related to operational excellence and politics. I expect this will improve with AI adoption and integration. You can't get an exec to create a decision record and commit it to git. Managers have incentive to sequester information.
Engineering already has the discipline (maybe) and abilities to solve the problem. Version control, change control, ADRs, logging, structured docs, etc... We can trace an inbound packet or call through the entire stack. Management can't/won't do anything remotely close. 1-to-1 emails, meeting minutes, stale Word docs is the standard for most.
Inserting LLMs as the interface, and/or plugging into existing interfaces like email, is going to change things. Finally it will be possible to capture more institutional knowledge, without trying to teach an old dog new tricks.
georgeburdell 5 hours ago [-]
In my experience, management is mostly pissing away the gains made by AI by either:
1. Pursuing polish and quality beyond previous norms
2. Replacing $100/mo/seat SAAS with something coded by a junior costing $200/day to develop over months.
The cost of code approaches zero, but the cost of having accountability, and hosting remains the same, and so individuals need to only coordinate to the extent that those things remain finite resources. Management needs to stop insisting that their directs adopt each others vibe coded tooling.
sanderjd 5 hours ago [-]
I think #1 is quite a good thing, on net.
650 3 hours ago [-]
"Shield the team from the business." The assumption underneath this one is that attention is finite and context switching is expensive. That assumption is intact. What changed is the cost of starving the team of context. Engineers prompting AI tools without business context just produce fluent, plausible, wrong work, at scale."
Too many teams and organizations have business types, mostly PM's who seek to lord over their area of know how and see themselves as delegators and mini CEOs, actively avoid looping engineers in to validate themselves. Engineers need to take on PM roles, and the PM role needs to be 1:50+ eng or go.
Mohiuddin7 53 minutes ago [-]
i agree with that but now at the pace where everything is moving really really fast, humans are using ai tools n all for imensively writing code and building new feature where they are also facing issues and frustration while debugging...llms are really good but sometimes they catch the issues but majority of time, they just grep or bash the things and eventually ending up not losing context....but if we go with your proposed way, then if humans build n llms review then it will be the slow process.....Hope I put my point correctly
jboss10 5 hours ago [-]
> Gemini 4 helped with the editing.
Does this guy have access to Gemini 4 already?
I'm guessing Gemma 4 was happy to be mistaken for Gemini and didn't catch this mistake.
trollbridge 5 hours ago [-]
“Helping” is doing some heavy lifting in that sentence! It appears to be 100% AI.
ilovefood 5 hours ago [-]
Yes correct, Gemma. Will correct it shortly.
dataplumb3r 5 hours ago [-]
If you wrote an initial draft you should consider using a better model or just posting what you wrote.
The smaller models can be sufficient for coding but for document writing not highly specific I've yet to be satisfied with AI output. I certainly wouldn't expect gemma to produce good outputs.
ilovefood 5 hours ago [-]
Makes sense, perhaps I'll post the raw notes + more material with the next article, prior to any edit. I like to use local models, even if they are not that powerful precisely because of that, I'm forced to put in more thought and it's my signed off article in the end.
Thank you for the suggestion.
add-sub-mul-div 5 hours ago [-]
You don't build credibility by putting out slop and correcting it every time someone points out that it's wrong. You think it's our job to proofread and fact-check?
ilovefood 5 hours ago [-]
A genuine mistake. I make these posts for myself, mainly to structure & share my thoughts. Let me know if you find other errors, I appreciate the feedback.
siliconc0w 5 hours ago [-]
I don't think it changes much for good managers. It should always be able setting people and processes up so the team can land durable measurable impact. The managers that thought the job of software engineers was to write code were bad managers. PRs or LoC were never good metrics.
ilovefood 5 hours ago [-]
Agreed, which is why I also believe "token usage" and similar proxy metrics aren't the right ones for what's next.
1over137 1 hours ago [-]
The cost of code has collapsed? Have programmer salaries gone down?
trollbridge 5 hours ago [-]
Pangram reports this post was 100% AI generated.
CharlesW 4 hours ago [-]
Pangram's marketing always reminds me of Anchorman's Sex Panther cologne: "They've done studies, you know. Sixty percent of the time, it works every time."
Did you read the paper you are linking to? It records a 0% false positive rate for evaluation of human-authored controls, and fewer than 5% of the hybrid and humanized papers had their AI levels overestimated by pangram. It is not completely clear from the data, but it seems that depending on whether the n=2 overestimates were “100% ai generated” assessments, the study you link says pangram’s 100% AI assessments were right anywhere from 95-100%
generic92034 49 minutes ago [-]
Still, I guess if you would care you could let the LLM perform loops with pangram tests, changing the text till it scores well enough.
2 hours ago [-]
CharlesW 1 hours ago [-]
My citation was correct. To question to ask yourself: For my use case, is it okay that Pangram can't reliably tell the difference between "100% AI generated" and "AI assisted"?
2 hours ago [-]
bonzini 5 hours ago [-]
It's depressingly hard to find one that isn't. You see the title, think this might be interesting, and puke by the second paragraph.
ilovefood 5 hours ago [-]
I'm really open to feedback. I checked your past comments and you posted:
> AI is good at coding if there's an oracle. If the system is ancient, unreadable, untestable, that's exactly the opposite. It won't get the exact set of corner cases.
I sort of mention this in the article, so I'm sure we're somewhat aligned on the core. How would you have worded things?
1 hours ago [-]
bonzini 1 hours ago [-]
First of all I want to say it is not your fault, and also your article is clearly partly human, despite what Pangram says.
The thing that I like the least is the headings and the way LLMs always try to put a punchy line. "Correctness time splits in two" says nothing if you don't know what it splits in. Maybe "making it correct vs. describing what's correct"?
Another trope is short sentences: "Good engineers used to say this before LLMs, and now nobody can argue about the sunken cost of having written that code." instead of the longer and redundant "Good engineers said this before LLMs. It was true then. It is enforceable now in a way it was not, because nobody can argue that writing more code was the hard part".
Another clearly AI paragraph is "Plumbing time collapsed. Scaffolding a service, generating tests, translating between frameworks, writing the first draft of a migration: all of this is fast now, and any timeline built on those costs deserves compression." Instead: "The time to bring up a proof of concept or refactor old code has compressed, and you should take that into account when planning your timeline".
There are videos on YouTube about AI style, you just need to learn them and undo them when they're the most blatant.
tiago_human 5 hours ago [-]
AI reduces the cost of writing code, but it makes adding things that nobody uses even cheaper. The bottleneck then becomes deciding what deserves to exist—and having the discipline to remove the rest. So this will imply more time spent in code reviews that lead to more iterations in PR's.
AI is doing a good job on writting code these days! Nothing against it; I use it every day, but the context switching is costing us a lot!
raffraffraff 5 hours ago [-]
My take, as a non-coder (well, not software engineering, I write 'code' but it's infra, and utilities in go/bash/pythong)...
I work at a company where the biggest problems are not 'writing code', they are:
- Organising teams
- Designing the system
- Prioritisation of work
The fuckups that we make on a daily bases are not 'code errors' they are failures in THOSE three things. I'll go into detail if anyone cares.
treetalker 4 hours ago [-]
Generative language models have helped me most by drawing my attention to the importance of context, communicative compression, prompting, comprehension, coherence, and coordination in the domain of human groups.
sanderjd 4 hours ago [-]
Same as it ever was.
convolvatron 5 hours ago [-]
don't forget 'actually making decisions'
whinvik 3 hours ago [-]
Sorry, "Gemini 4 helped with editing"?
0gs 3 hours ago [-]
dang it i tried to check for a dupe but there were too many comments!
antonvs 6 hours ago [-]
The worse problem is blog posts after the cost of writing collapsed.
Not everything has to be written as though it’s a middle manager’s idea of what makes for a good TED talk.
nvme0n1p1 5 hours ago [-]
I closed the tab after seeing the AI hero image. Good to know I didn't miss anything.
JSR_FDED 5 hours ago [-]
You did miss something. But the picture doesn’t help…why would you start an article with an image that telegraphs “low effort”?
zo1 5 hours ago [-]
Reading the domain-name had the same effect for me.
obmelvin 5 hours ago [-]
You discount people's blogs based on them having a non-western name? Maybe I've misunderstood your comment, but the domain is just their name.
ilovefood 5 hours ago [-]
I can't change my name, but happy to change the hero picture :))
aplummer 5 hours ago [-]
I read this entire article and didn’t sniff AI, plus it had some good insights…?
nwah1 5 hours ago [-]
Feels like a human-authored outline that was run through AI to make it into an article.
ilovefood 5 hours ago [-]
I've written this myself, Gemini 4 did a bit of editing:
> What follows is a cleaned-up version of notes I accumulated over the past year. Gemini 4 helped with the editing.
The image is made by an AI image generator on fal.ai. It's better I spare you all my design skills :)
hgomersall 5 hours ago [-]
Well it contains a typo, so it made a cock up.
lcampbell 5 hours ago [-]
More than that, the AI disclaimer itself is a hallucination:
> Gemini 4 helped with the editing.
There is no Gemini 4, unless the author is writing from the future.
ilovefood 5 hours ago [-]
You are correct, my bad. I meant to write Gemma 4. I'll update.
sodapopcan 5 hours ago [-]
AI image, though.
JSR_FDED 5 hours ago [-]
Yes the article took a while to get going, but once it did it was thoughtful and well reasoned.
ilovefood 5 hours ago [-]
Thank you very much for the kind words :)
cineticdaffodil 6 hours ago [-]
Can a llm predict the price of a change to a codebase in tokens and predict the origin of the price, aka cam it see good and bad architecture?
arjie 6 hours ago [-]
Some code bases better than others but the top models can.
coffeebeqn 4 hours ago [-]
Sure and having some in repo documentation markdown makes it a lot easier
mathgeek 6 hours ago [-]
Ask it to do so, would love to know your results.
cineticdaffodil 5 hours ago [-]
I did, the problem is that you need a sort of standardized feature change to compare similar repos. So the metric is relative only to same projects lacking that same feature. So no, its a useless thing. No architecture score comparisson between apples and oranges.
Software development has always evolved. Sometimes slowly, sometimes quicker.
LLMs have brought a different unlock, and for everything we're seeing become easier, it allows people learn to use the tools to take on solving problems that couldn't be approached before.
glimshe 4 hours ago [-]
Your short post speaks of exactly what people should be talking about. Due to the panic related to job replacement, not many are talking about the new frontiers. Vibe coding is boring because it does the same faster/cheaper. I'm more interested in the things we couldn't do before but we now can because there's a crazy savant a few keystrokes away.
happytoexplain 5 hours ago [-]
Code was never expensive.
doug_durham 5 hours ago [-]
Code was always expensive. Entire industries of design tools cropped up simply to avoid writing the wrong code since it is so expensive. There are entire journals dedicated to this. Organizations are fixated on making sure that the right code is written since the cost of getting that wrong is so great.
SoftTalker 4 hours ago [-]
Building the wrong thing is what is expensive.
Once you know what you are building and can clearly describe it, the code isn't the hard part. Or at least that's how it has always seemed to me.
mathgeek 5 hours ago [-]
I assume GP meant exactly this (the people designing and writing good code are expensive).
mohamedkoubaa 4 hours ago [-]
At least not the writing of it
jdw64 2 hours ago [-]
Coding used to be an expensive task. Seeing how quickly repositories have grown with the rise of AI coding, it's clear how much people wanted to build things but were thirsting for the means to do so. Even languages like R, which were mostly used by graduate students and experts, have seen a massive increase in usage since vibe coding became popular.
Honestly, when people say AI code quality is bad, Linus himself has said it's now genuinely useful. AI is useful and writes better code than most people. Even in competitive coding, tourist lost to AI. And in the most logical field of all, mathematics, AI is churning out an enormous number of theorems.
Looking at all this, it's fair to say AI is at least at a PhD level of technical ability, and most people would admit they don't have PhD level skills. Of course, there are still many people who code better than AI. But at least when it comes to unfolding logical structures, AI has a higher chance of being more logical than humans. Within a given framework, AI constructs much more logical structures.
That's why I think the article's use of the word 'semantic' is right. It's humans who form the framework, and that's the semantic, while AI fills the empty spaces inside it. If you feed it a flawed framework, it fails.
And the fact that AI is more logical than humans is paradoxically a greater risk. Human developers can rely on tacit knowledge to make reasonable compromises even when the requirements, the framework, are sloppy. AI can't do that. If there's a logical gap in the framework humans design, AI will exploit that weakness and expand the state space into regions we can't cognitively grasp.
Programming is ultimately about how you occupy state space. The problem is that as the program grows, the cognitively inaccessible territory keeps expanding. So we distribute trust across reliable points, libraries, frameworks, and for my own code, once it exceeds tens of thousands of lines, I rely on tests and gates.
Honestly, the idea of understanding everything in a program is a purely academic claim. Once the program gets large, it's impossible. No one can know every external factor, test bug, or unexpected interaction.
The issue is that with LLMs, when the prompt input goes deeper into the semantic space, it also reaches into areas I don't understand, producing code at a depth that's untestable.
For example, I might be an expert in domain A but a beginner in domain B. If I inject expert level knowledge for domain A into the AI, the AI will try to match that level in domain B as well. That results in code I can't understand or modify, and eventually, I'm left with no choice but to replace all the code with AI generated code.
So I'm wondering what to do about this. Should I focus on gaining empirical experience in handling black boxes? Or should I stick with smaller, human written codebases?
But realistically, the current situation, where I can build bigger and touch more things, is more enjoyable to me. I think what I actually enjoyed wasn't programming itself, but the act of creating something.
chasd00 49 minutes ago [-]
> I think what I actually enjoyed wasn't programming itself, but the act of creating something.
i think this is the real difference between the pro-ai and anti-ai crowds.
OutOfHere 3 hours ago [-]
Engineers do not need management. Investors do.
As I understand it, the purpose of management is to match financial resources with material+human resources to perform feasible tasks. There is nothing here I see that can't be done by an experienced token generator. If anything, automating management seems easier than automating engineering.
As for leadership, it can be done by the investors.
visarga 2 hours ago [-]
You can automate work but never owning the consequences. AI tasks emerge from human contexts, they perform work in the context and finally the outcomes collect in the context - gains, losses, risks, costs. So AI is great but it needs our skin for the start, middle and end of a task.
OutOfHere 2 hours ago [-]
It's not as if management owns any consequences. At best they adapt, but so can AI. Management tasks are not like engineering tasks.
pillefitz 50 minutes ago [-]
As stated in the sibling comment, management is very similar to engineering.
pillefitz 52 minutes ago [-]
This is such a simplistic view I have difficulties deciding where to even start. Think of a manager as someone who engineers the fabric of the organization. Most of my time is spent debugging, communicating, finding the right abstractions, configuring processes - all very similar to the engineering work I previously did.
chrisjj 4 hours ago [-]
> The cost of producing plausible code has collapsed
Who wants merely plausible code?
devin 2 hours ago [-]
Your director who is burning tokens making a POC you now need to figure out how to support at scale.
matthorse 1 hours ago [-]
[dead]
2596-ANXC 3 hours ago [-]
[dead]
vips7L 1 hours ago [-]
I miss when this forum used to talk about computer science and programming languages.
phrones1s 5 hours ago [-]
[flagged]
CurbStomper 6 hours ago [-]
[dead]
lardosaurusrex 5 hours ago [-]
[dead]
stefangordon 4 hours ago [-]
I’m struggling to imagine a project that would require more than one talented engineer and a bucket of tokens anymore.
I think perhaps the assumption that engineering managers should have any employees may be outdated.
I can imagine average and mediocre engineers equipped with tokens could create chaos and debt on a scale never before imaginable, so it’s easy to see how orgs who still have these employees around are struggling with the transition.
The reality is you need to get rid of them all, and replace them with the most experienced highest paid person you can find. In the near future that person will become obsolete too.
sdevonoes 19 minutes ago [-]
Would that single person handle different topics such as: UI/UX, security, backups, distributed systems, data migrations, infrastructure, …?
Sure thing there are engineers out there that know some about all of the above (I personally do all of that on personal projects) but you still need specialists, otherwise it’s you alone with all the unknown unknowns that the llm may claim to solve, but you cannot verify
chasd00 54 minutes ago [-]
> I think perhaps the assumption that engineering managers should have any employees may be outdated.
to me, the latest models and harnesses turn all devs into an engineering manager with one direct report. Some devs naturally take up the role and great things happen, for others it's like trying to get a fish to ride a bicycle. This is how it was pre-genai too so i don't think there's a right or wrong answer, some will take off with the technology and some will struggle.
chrisjj 3 hours ago [-]
> I’m struggling to imagine a project that would require more than one talented engineer and a bucket of tokens anymore
The assumption is that LLMs should be writing the code and human engineers reviewing and verifying the LLM output. And that this pushes the cost of producing down. And I fundamentally disagree with that.
Every time I ask LLMs to write code, even with Opus 4.8 (haven't tried it with Opus 5 yet), what I get ends up being totally rewritten. LLMs still aren't good at writing maintainable code. Can they write plausibly functional code? Yes. But it won't survive the long term. People using LLMs to write all their code are gambling on them eventually getting to a point where the LLMs can fix their own code. It's possible, but I wouldn't necessarily bet on it.
Where I have found immense value from LLMs is in code review. Repeated review by LLMs catches an amazing amount of potential issues. They really shine on security review, but are very effective with any kind of review.
The other thing that the "LLMs write code camp" misunderstands is that writing was never the bottleneck. Understanding was. And understanding the code is still the bottleneck. But understanding is truly gained during the writing loop. The understanding you gain from pure reading or code review is marginal compared to the understanding you gain while writing.
Most of the time previously spent writing was actually spent updating and deepening our understanding of the system under development. There's no replacement for that understanding in a world where LLMs are doing the writing.
But if you flip it: humans write, LLMs review, then you still get a major gain -- not in speed, but in quality. And you keep the understanding loop intact. I would propose that this might be the best way to deploy LLMs.
But, with AI, the cost of code has gone down more than it already has. Well, cheap things are easy to throw away. So you prototype, prototype, prototype, and close the loop as much as possible with the customer. True agile development, not big A Agile.
The problem is this requires alignment from management, and we're just not seeing it at many company. They can't grasp that things have changed, and that throwing away code is free. They don't trust engineers to close that gap, so customers and stakeholders are still waaaaaay over there and we're delivering features they don't want.
Until managers give up on Agile, we'll never get time to actually write specs
To be honest, the models are getting so good that they do most of this unprompted now.
At this point if you can’t get the agent to write good code then either I) you are in a very specific niche (like Karpathy trying to write NanoGPT) that is extremely out-of-distribution, or II) skill issue, you need to learn how to prompt better.
It’s fine to have a skill gap! Just don’t delude yourself that the tools are bad and everyone claiming they are good is wrong.
So despite its importance much of it is actually pretty in-distribution.
Well put. This insight is worth repeating in every discussion on the subject, from software engineering to mathematics.
The problem is that understanding is not the product being sold. The business model is for everyone to become consumers of what the magical genie generates, where the "understanding" is kept on the side of the model providers. This ensures a future generation of consumers dependent on someone else to provide the understanding.
Otherwise, you can create your own answers based on actually understanding the code, theorem, proofs, etc. Smart consumers of LLMs will use them to increase their own knowledge and understanding, so that the service is augmenting their intelligence, not replacing it.
But to be fair human code reviews have the same problem. It's like reviewers feel they have not done their job if they don't find something wrong.
I do the same for LLM code review comments: some changes are out of scope or could be moved to a separate PR; some edge cases don't happen in practice and should just fail noisily instead of writing more code to maintain. When these are the only issues it's raising, then I know it's done.
I wonder if the same trick works with AI?
Perhaps someone more knowledgable could jump in here to clarify?
To be clear I mostly only use Opus and Gemini Flash but this might work for others too.
Disabling your C compiler warnings works too. You get to ship then leave work early!
The key I’ve found is human peer review. The reviewer jumps on a live call with the developer, pulls up the PR with transcription on, and asks questions. At the end of the call, the transcript passes back into the coding agent and the PR is polished up, becoming more self-documenting, and the humans are left with some degree of common understanding of what’s going on.
I’ve been operating my team of ~15 this way for 9mo to great effect… there is simply no going back to the stone ages.
Can I watch/observe one of your review sessions?
Every few weeks, I hear the beginnings of a great approach towards working with LLMs but I rarely see it in practice.
If you're open to this, remote or in person, ping my username at gmail.
But that doesn't mean they're useless either. I use them all the time for review as you mentioned or to knock out one-off scripts that don't go anywhere near source control. There's just no world where I don't need to understand every line of code that I'm responsible for getting into our project.
If LLMs are, as you stated, really good at catching potential issues, then they are, almost by definition, really good at producing code without potential issues, if guided correctly: all they need to do is inspect and iterate, until they do not find any more potential issue in the code they produced.
This was my stance a couple of years ago, but now I've given it up.
It turns out writing actually was the bottleneck. You can understand perfectly well what you want, but writing it is long and tedious to the point where you find excuses not to do it. Particularly with version 2, the step where you have an OK system and you want to improve it. Quite a lot of changing the code is just useless busywork: re-wiring old functions, moving imports around, searching for locations that benefit from extracting a common piece of code. And each time you do one of those, there's a decent chance you did something even more trivial like forgetting a semicolon or calling the wrong function.
Now that I have an LLM helping me, I can see why. The critical decision is a terse declarative like "we need to have several TCP connections instead of one, and just use the sequence number to arbitrate". A human junior programmer could perfectly well understand what this meant, but he would have to go through all of the above to get to the final product. Now, I can just tell the LLM and I will get what I want, even with the things I didn't explicitly state, without spending attention.
This means I can use my attention on the things that matter. So instead of spending today thinking about how to arbitrate between the TCP connections and tomorrow thinking about pre-calculating my outgoing orders, I can just do both today. I don't waste the good waking hours chasing minor bugs, I just think about the large structure.
I get the feeling the best programmers of years past were actually masters of the little things, which led them to be able to look at the big things. Essentially it was cheaper for them to get to the top of the mountain, where you can see the landscape. Kinda like how the kid who was good at mental arithmetic in primary school was also good at calculus at the end of high school: if you don't have to concentrate on the little things, you have time for the big things.
how did you decide to pick the most trivial kind regression for this example? do you compile your code before checking it in?
> A human junior programmer could perfectly well understand what this meant, but he would have to go through all of the above to get to the final product. Now, I can just tell the LLM and I will get what I want, even with the things I didn't explicitly state, without spending attention.
the main efficiency you have described here is offloading the verification of a change onto the LLM. that is the bottleneck. readers can decide whether a non-deterministic statistical model is a good tool for this job
> I get the feeling the best programmers of years past were actually masters of the little things, which led them to be able to look at the big things
the best programmers understand that their job is to automate workflows, and that includes their own. if you're worried about missing a semicolon, I'm sorry to say that's a skill issue
Why would this be a regression? You might just be writing a new line of code.
> do you compile your code before checking it in?
Well obviously. That is generally how you discover that a semicolon is missing.
> the main efficiency you have described here is offloading the verification of a change onto the LLM. that is the bottleneck. readers can decide whether a non-deterministic statistical model is a good tool for this job
No, it's the time between you deciding something needs to be done, and it being done, that is the bottleneck. You cannot avoid trying to compile the code and testing it. Now you can get to that test without paying attention, which is time you can use productively.
> readers can decide whether a non-deterministic statistical model is a good tool for this job
Somehow, the non-deterministic model has built me the deterministic code that I want, very fast, pretty much all the time. A year ago it would get stuck. Now it doesn't, for me at least, and for competent programmers that I know.
> the best programmers understand that their job is to automate workflows, and that includes their own. if you're worried about missing a semicolon, I'm sorry to say that's a skill issue
Well yeah, and I've automated my workflows completely. I don't have the problems I used to have. If you haven't caught on to the new way of working, well, that's a skill issue...
I still find the models get stuck or go on _massive_ side quests. Just today, I asked claude to write a hello world C++ program using import std; I interrupted it when It decided I needed a new toolchain installed, and started checking for docker installations. This is super basic stuff, it hadn't even generated a plan, it just started searching for LLVM versions rather than running clang --version.
> If you haven't caught on to the new way of working, well, that's a skill issue...
Honestly, it feels like the emperor has no clothes on this topic, and the crowd defending LLMs to death are way too quick to call it a skill issue.
I feel it's the other way around. The LLM skeptics are unwilling to admit that these things can get you there faster than you would on your own, in the face of clear evidence.
I use LLMs, but they're just a tool in the workflow, and I make sure to review the output. they might remember semicolons but they make much more pernicious mistakes that are harder to detect
If other people are dissatisfied with LLM output quality while it seems to work fine for you, you might want to consider that the quality of code you produce is closer to the quality of code the LLM produces than what those other people are producing.
What you posted there, for example, about most of changing code being busy work is a pretty big red flag for a codebase. One of those "large structure" things that you're supposed to be paying attention to is the architecture of the code. There's always the chance that some change you need to do goes against the grain of the solution you architected, and you need to make changes all across your codebase to fit it in, but in general the point of modularity and good architecture is that when you make a change you just have to make that one change, ideally just changing the logic of the one responsible function with only minor changes required anywhere else in the codebase. If you're consistently having to hunt throughout the code for related functions that you need to rewire that's a sign that your architecture does not fit with the direction your codebase is evolving, or alternatively that you don't have much of an architecture to begin with and your code is highly interconnected.
Actually one habit you mention at the end of that quote can worsen this issue: "searching for locations that benefit from extracting a common piece of code". Tautologically this is a good thing as you define it as only working on locations that will benefit, but given the frequent need for rewiring of functions I would hazard to guess that you've "deduplicated" code a bit overzealously. Just because two functions share some common code does not necessarily mean it is appropriate to pull that out into a function. Deduplicating is good if conceptually the code is a single thing that you would always want to keep in sync, as it means that when you need to make a change to it you don't have to hunt down all the places it's used. On the other hand, if you find yourself frequently needing to delve in to these functions to rework them because you need to make a change to how it's used by just one caller, your "deduplication" has added to your workload, and probably created some overcomplicated code in the function that is in reality handling multiple distinct needs.
I hope this doesn't come across as too condescending, and if I've just wasted your time explaining principles you already understand I apologize. I don't know you or the code you're working on so I can't exactly confidently judge your work solely on a few paragraphs. It's just that your mention of how your experience of coding has been different from what others have described, and specifically that, for you, writing has been the bottleneck rather than understanding, combined with the specific issues you describe facing, imply to me that you may not realize that the approach you are taking to producing code yourself may be significantly different from how other Software Engineers are producing code, and that may account for some of the differences you note in your personal experiences programming.
1) It was good for me to spend years learning the little stuff. Loops, variables, if conditions, how to import stuff, git, debugging things, reasoning about the flow of control. Classic coding.
2) I had a false dawn at about 10 years in. I thought I understood a lot.
3) I learned I had a lot to learn. Very wide areas of programming I'd never touched, ways of thinking that started to click.
4) I spent another ten years covering holes, building a different type of experience. My guesses about how to do a project are much better now. My guesses about what really matters have changed.
5) Now the small stuff is actually just bothering me. I'm not going to learn much more from staring at little things. There are larger architectural things to think about, and the little things are just friction.
So that's where I'm coming from. I get that a lot of pushback is going to be from 10-year-me, who thought he'd gotten to a high level of understanding by slogging through the little stuff.
Speak for yourself. Writing code has never been a bottleneck for some of us. I can't speak for everyone, and neither should you.
>Now, I can just tell the LLM and I will get what I want, even with the things I didn't explicitly state, without spending attention.
This should worry you. All too often the LLM invents things I didn't ask for and implements things I didn't need. YMMV, I guess. If slop gets the job done, and nobody notices, then who should care?
Are they as good as handcrafted code by 0.1% of top software engineers. Generally no. But neither is 99.9% of real code.
LLMs also are good at code reviews. What they'll miss is often the big picture but they can still catch plenty of issues. I still want to see a human in the loop in my domain.
Totally agree that writing the code was never the bottleneck. We're not seeing massive productivity gains even if some code is written faster. It's not just about understanding but also various other activities that happen in large companies and teams.
Also agree LLMs can be used to gain quality but realistically most orgs are going to aim for "fixed or decreasing" quality at lower costs.
However, sometimes then I tell it to write an app with detailed instructions and it spits out garbage so your mileage might vary.
Debugging code step by step is how I understand complicated code.
I do know what good code looks like, but does that even matter anymore? All I know is that now, I get to focus on endless UX polish, which is the only thing the matters.
I feel like we are living through something like the Protestant Reformation, where priests once spoke Latin, and then started to speak in plain local language. The old guard did not like this.
I've been using it for a unity game for the past few years. Nowadays it will go sleuthing into packages and assembly and make decisions based upon what it sees there.
It will make comments about why it's doing something based upon a function call 3 methods deep.
God forbid any of these details change in a minor version update.
I find that depends on the target language. They can be good at writing maintainable code, but not consistently across every language.
The languages beginners usually gravitate towards are especially hard for LLMs to produce quality output for. Presumably this is due to the training data including all the unmaintainable codebases written by beginners in those languages, which hasn't allowed the LLM to converge on recognizing what a maintainable codebase looks like in those languages.
If programmer productivity was something we actively optimized for, we wouldn't have crammed programmers like sardines in warm and noisy open floor offices with 2000 ppm CO2 levels and then further constantly interrupt them with emails and slack pings and meetings all day long, Jira rigmarole wouldn't make up a significant portion of what they did, programmers would have instead mostly been thinking and programming.
We've always had the ability to 2X if not 10X the output of each and every one of those poor souls. You don't end with this sort of programming purgatory because it's a productivity optimum, it very clearly isn't, but because it's a billable hours optimum and/or an org chart clout optimum and/or because of Jevons paradox got hands even in business management and the IT department was allocated too many dollars.
I disagree. Kinda
What AI has made much simpler is that you don't have to waste time checking docs and have the best autocomplete system by a long shot - this was a bottleneck unless you were doing Java or some other language with "perfect" AC
What AI made "kinda easier": solving for usual problems. The stuff you would search Stack Overflow, or think a couple of minutes for an optimized solution - not a bottleneck but not 100% smooth neither
You still have to test and validate your code. AI made this easier-ish but this is still where I see manual work being needed (even if you are automating tests - you still have to think on what you want the code to do)
LLMs help, but they haven’t been trained on our own repos. I don’t need the LLM to help me with algos that are available online… I need them to help me with custom business logic
I think now, code is the bottleneck. Just because you can generate million lines of code, people with different skill level think they are accomplishing the task, testing, merge conflicts, trust has become the bottleneck.
The “old way” would be lots of debate (both bike shedding and useful) among engineers during design phase, and then you’d implement.
Now it’s shifted so there are no design docs and there is only the generated prototype. People trying to do their design review while there’s already a functional-ish prototype and it goes nowhere. There’s an anchoring effect in place because the first thing already exists and management says “this seems to work, just use it and move on”. The result is that useful debates about substantive issues don’t happen and bikeshedding is all way get to do
And then you get paged at 2am because prod is down and the support channel is more active than the team's one.
In larger organizations, quite often it's the business that is holding back development. They can only handle so much change and speed needs direction to be velocity. Drafting requirements is generally much slower than implementing them.
Like the number one complaint from programmers has been that they don't get to do programming. They want to write code, not update jiras or spend hours in meetings.
There are a lot of (excruciatingly) long-form posts about what folks are pioneering but not a whole lot of follow up about what failed. Where are the short posts on the negative space? How did halving your staff work out? Flattening your org? All those dark factories, what haven't they produced? How about all the other things tried, failed, and unceremoniously scrapped?
We need to explore and communicate the negative space more efficiently. Don't repeat the same mistakes, and don't make me read 2653 words when 300 do it better.
Similar to a functioning side project in the 5-10k LOC range. Announcing something that worked a year ago, was laudable, even if not profitable.
I vibe coded 15k LOC this morning and read 20k words of AI generated text while doing so. No longer are either noteworthy or valuable public contributions just by virtue of having been done. I don't think that's widely recognized yet.
Isn't that the essence of "attention is all you need"?
A million times this. Can we please RL the next models to learn the “if I had more time I would’ve written a shorter letter” method please.
I see it every day in tickets, many communication channels, PR descriptions, comments, documentation. All have at least 70% verbose fluff which is so taxing and makes it very hard to keep track of the one important thing they’re trying to communicate in the message
New tech stacks will appear and they will handle the mess to some extent.
This is a long-standing problem related to operational excellence and politics. I expect this will improve with AI adoption and integration. You can't get an exec to create a decision record and commit it to git. Managers have incentive to sequester information.
Engineering already has the discipline (maybe) and abilities to solve the problem. Version control, change control, ADRs, logging, structured docs, etc... We can trace an inbound packet or call through the entire stack. Management can't/won't do anything remotely close. 1-to-1 emails, meeting minutes, stale Word docs is the standard for most.
Inserting LLMs as the interface, and/or plugging into existing interfaces like email, is going to change things. Finally it will be possible to capture more institutional knowledge, without trying to teach an old dog new tricks.
1. Pursuing polish and quality beyond previous norms
2. Replacing $100/mo/seat SAAS with something coded by a junior costing $200/day to develop over months.
The cost of code approaches zero, but the cost of having accountability, and hosting remains the same, and so individuals need to only coordinate to the extent that those things remain finite resources. Management needs to stop insisting that their directs adopt each others vibe coded tooling.
Too many teams and organizations have business types, mostly PM's who seek to lord over their area of know how and see themselves as delegators and mini CEOs, actively avoid looping engineers in to validate themselves. Engineers need to take on PM roles, and the PM role needs to be 1:50+ eng or go.
Does this guy have access to Gemini 4 already?
I'm guessing Gemma 4 was happy to be mistaken for Gemini and didn't catch this mistake.
The smaller models can be sufficient for coding but for document writing not highly specific I've yet to be satisfied with AI output. I certainly wouldn't expect gemma to produce good outputs.
Thank you for the suggestion.
Pangram's "100% AI generated" claims are right 65% of the time. https://link.springer.com/article/10.1007/s40979-026-00226-w
> AI is good at coding if there's an oracle. If the system is ancient, unreadable, untestable, that's exactly the opposite. It won't get the exact set of corner cases.
I sort of mention this in the article, so I'm sure we're somewhat aligned on the core. How would you have worded things?
The thing that I like the least is the headings and the way LLMs always try to put a punchy line. "Correctness time splits in two" says nothing if you don't know what it splits in. Maybe "making it correct vs. describing what's correct"?
Another trope is short sentences: "Good engineers used to say this before LLMs, and now nobody can argue about the sunken cost of having written that code." instead of the longer and redundant "Good engineers said this before LLMs. It was true then. It is enforceable now in a way it was not, because nobody can argue that writing more code was the hard part".
Another clearly AI paragraph is "Plumbing time collapsed. Scaffolding a service, generating tests, translating between frameworks, writing the first draft of a migration: all of this is fast now, and any timeline built on those costs deserves compression." Instead: "The time to bring up a proof of concept or refactor old code has compressed, and you should take that into account when planning your timeline".
There are videos on YouTube about AI style, you just need to learn them and undo them when they're the most blatant.
AI is doing a good job on writting code these days! Nothing against it; I use it every day, but the context switching is costing us a lot!
I work at a company where the biggest problems are not 'writing code', they are:
- Organising teams
- Designing the system
- Prioritisation of work
The fuckups that we make on a daily bases are not 'code errors' they are failures in THOSE three things. I'll go into detail if anyone cares.
Not everything has to be written as though it’s a middle manager’s idea of what makes for a good TED talk.
> What follows is a cleaned-up version of notes I accumulated over the past year. Gemini 4 helped with the editing.
The image is made by an AI image generator on fal.ai. It's better I spare you all my design skills :)
> Gemini 4 helped with the editing.
There is no Gemini 4, unless the author is writing from the future.
LLMs have brought a different unlock, and for everything we're seeing become easier, it allows people learn to use the tools to take on solving problems that couldn't be approached before.
Once you know what you are building and can clearly describe it, the code isn't the hard part. Or at least that's how it has always seemed to me.
Honestly, when people say AI code quality is bad, Linus himself has said it's now genuinely useful. AI is useful and writes better code than most people. Even in competitive coding, tourist lost to AI. And in the most logical field of all, mathematics, AI is churning out an enormous number of theorems.
Looking at all this, it's fair to say AI is at least at a PhD level of technical ability, and most people would admit they don't have PhD level skills. Of course, there are still many people who code better than AI. But at least when it comes to unfolding logical structures, AI has a higher chance of being more logical than humans. Within a given framework, AI constructs much more logical structures.
That's why I think the article's use of the word 'semantic' is right. It's humans who form the framework, and that's the semantic, while AI fills the empty spaces inside it. If you feed it a flawed framework, it fails.
And the fact that AI is more logical than humans is paradoxically a greater risk. Human developers can rely on tacit knowledge to make reasonable compromises even when the requirements, the framework, are sloppy. AI can't do that. If there's a logical gap in the framework humans design, AI will exploit that weakness and expand the state space into regions we can't cognitively grasp.
Programming is ultimately about how you occupy state space. The problem is that as the program grows, the cognitively inaccessible territory keeps expanding. So we distribute trust across reliable points, libraries, frameworks, and for my own code, once it exceeds tens of thousands of lines, I rely on tests and gates.
Honestly, the idea of understanding everything in a program is a purely academic claim. Once the program gets large, it's impossible. No one can know every external factor, test bug, or unexpected interaction.
The issue is that with LLMs, when the prompt input goes deeper into the semantic space, it also reaches into areas I don't understand, producing code at a depth that's untestable.
For example, I might be an expert in domain A but a beginner in domain B. If I inject expert level knowledge for domain A into the AI, the AI will try to match that level in domain B as well. That results in code I can't understand or modify, and eventually, I'm left with no choice but to replace all the code with AI generated code.
So I'm wondering what to do about this. Should I focus on gaining empirical experience in handling black boxes? Or should I stick with smaller, human written codebases?
But realistically, the current situation, where I can build bigger and touch more things, is more enjoyable to me. I think what I actually enjoyed wasn't programming itself, but the act of creating something.
i think this is the real difference between the pro-ai and anti-ai crowds.
As I understand it, the purpose of management is to match financial resources with material+human resources to perform feasible tasks. There is nothing here I see that can't be done by an experienced token generator. If anything, automating management seems easier than automating engineering.
As for leadership, it can be done by the investors.
Who wants merely plausible code?
I think perhaps the assumption that engineering managers should have any employees may be outdated.
I can imagine average and mediocre engineers equipped with tokens could create chaos and debt on a scale never before imaginable, so it’s easy to see how orgs who still have these employees around are struggling with the transition.
The reality is you need to get rid of them all, and replace them with the most experienced highest paid person you can find. In the near future that person will become obsolete too.
Sure thing there are engineers out there that know some about all of the above (I personally do all of that on personal projects) but you still need specialists, otherwise it’s you alone with all the unknown unknowns that the llm may claim to solve, but you cannot verify
to me, the latest models and harnesses turn all devs into an engineering manager with one direct report. Some devs naturally take up the role and great things happen, for others it's like trying to get a fish to ride a bicycle. This is how it was pre-genai too so i don't think there's a right or wrong answer, some will take off with the technology and some will struggle.
Try air traffic control?