Amplification, not Replacement — and Why I will NOT Frame AI as Developer Replacement

There is a story about AI that is easy to tell - and, oh so tempting to believe. AI writes the code now so you need less engineers and those left are more like watchers than creators.

It sells well. It ages badly. And I will be damned if that story gets told when the teams I lead. It's bad not just for your morale, but bad engineering.


The sequence that no-one puts on the slide

Let us assume that you truly take replacement story seriously and reduce developers to cut on AI. Trace the chain of events that we actually witness.

Tribal knowledge is the first thing to walk out the door. The context that lived in people's heads leaves with them: how this system behaves, what corner cases matter, and what assumptions are never written down.

Then the left over people end up owning huge portions of code they did not write and are unable to fully explain. Whilst the agents are still outputting more of it, at size and faster than anyone can begin to comprehend.

Finally a event touches an area of the system whose context just left the building. Because the understanding required to debug it is gone, the response is tepid. A problem that should be contained leads to an extended one.

You exchanged a salary line for a dark-code cascade. And in a year, it is more expensive than saved.

That is not a moral argument. It is an operational one.


A different framing

I tend to work from the opposite framing where people and AI together more than the sum of either alone.

So the engineer is not a supervisor over a machine that makes some things. They move up the stack. They write the specification. They investigate the output to see if it is correct. They take ownership for the behaviour in production. They determine what gets built and equally, perhaps more importantly, what does not.

Less typing you do. Your judgment is worth more.

And this is where the leverage is. Allegedly a model can spit out a reasonable implementation of near anything you describe. But it cannot tell you if the thing shouldn't be built in the first place, if this spec even represents what customers really need or not and whether XYZ edge case will hurt someone six months from now.

That is human work. And there is now more of this, not less.


Why the eyes are on leaders on this

The second reason for getting the framing right is about keeping the people you want to keep the most.

Your best engineers are reading the tea-leaves of what you say about AI in a moment of transition like this; Each is trying to determine if this is a compatible place for them moving forward.

The best ones are the first to leave because they are exactly the people who have options elsewhere, if what they hear is 'you're being phased out'. You lose the engineers you would have most liked not to lose and retain those who can no longer move.

If the message is: Your judgment becomes more important now, and the boring parts of being a job are going away, they lean in. They are the ones who ascend the learning curve faster than everyone else, and teach everybody else.

This framing is not spin, it is an accurate read of value shift. It also decides for whom in the room remains while you form that transitional shift.


Be honest about what changes

None of which means pretending nothing changes. People would know straight away that it is a fake and dishonest.

The day-to-day shifts. Other skills that were once core are no longer as important. New ones matter more:

Writing precise specifications.

Designing good evals.

Good rapid judgment of AI output.

To make that move people will need support. Be direct, and attach it to the direction which is right: more than the establishment, not replacement. The engineers are not being managed out of the way. They are being migrating to the region in the image where people cannot be replaced.

How are you framing AI for your people, and are you certain they hear the same message that you believe your sending.

How to Navigate Distributed Teams Through the Same AI Dip Without Re-Learning

For companies running engineering in more than one location, there's a risk to AI adoption that single site teams never face. This risk is not technical - it is organizational, and like any such risk it is easy to overlook until the price has been paid.

The risk is easily articulated - you end up paying multiple times for the same learning curve.


Parallel, uncoordinated adoption trap

Imagine an organization with three engineering sites. AI arrives. The measure is implemented at each site, composed of highly skilled personnel.

You are trained on a limited set of prompts and conventions over a site. One develops its own, distinctly conflicting process. One option is to wait and see what works with about one-third.

But each site hits the same dip, that trough right when you introduce an AI. The parties are processed through their own respective queues. They all learn the same lessons - the hard way - what should be delegated and what you must verify.

Six months later you look around and see three dialects of use of AI, three sets of standards half-built, and three teams who each paid the full price of the same mountainside. Nothing accumulated. You can't throw upward, so the effort was compounded sideways.

The technology was never the issue. The coordination was the problem.


Central control is not the fix

The obvious (and the incorrect) solution is to centralize. Select one approach, demand it everywhere, enforce consistency.

This fails for a very simple reason. Adoption does not happen like waste; it is rather the local experimentation that each site is doing. People learn about AI by testing it against their stack and under constraints specific to their organization. Impose one-size-fits-all working practices cross-border and you destroy the very thing that works about adoption.

Hence, the aim is not sameness of practice. That is, shared learning with local freedom.


What actually travels well

Three things make the difference between three teams climbing in isolation and one organization growing together.

Shared vocabulary. Everyone describes maturity and practice with the same vocabulary. One site can say a module is at the AI level and exact some other kaleidoscope knows with no translation. Then a conversation across sites is not a negotiation over terms, it is about the same thing.

Shared artifacts. All the configuration files, the prompt libraries, eval scenarios and security checklists live in one place and can move freely. One useful pattern discovered in one place is automatically available to all of them, rather than someone mentioning something during a call.

Local autonomy on top. Every site retains its own racers who scout on the ground, customizing the common foundation to their reality and broadcast outward to share their learnings on almost a monthly basis. They go local first, then report back across the borders.

It is not a top down architecture with central home base hierarchical dogs barking orders below (the shape), but an upper level light weight part of shared identity plus place-based exploration [above]. When it crystallizes, the organization moves once up the curve instead of every time for every site: the second site builds on first-site learning; and the third is built on a bigger base than any one started.


The sacrifice of a leader

There is a price to pay and it doesn't come so much with the teams as with the leader.

It means, of course, not giving in to the common inclination: letting every strong site operate at its own pace (because they can and probably would be okay without coordination). Instead you work at the connective tissue, the common index and cross-site sync and shared language.

This work has no demo. The only trace it leaves is that there was no pain, because the coordination worked so there is no more duplicated pain. And that is a tough thing to invest in, precisely because success looks like nothing happening. However, it makes a difference between AI making your entire organization better by using the tool across multiple sites vs AI tiring each of your sites separately.

If one of your sites learns something AI useful this month, how is that information shared with the others and what is the lag time?

The J-Curve Thing Nobody Tells You About: Why Good Teams Get Slow First with AI

The one that should be pinned on the wall for every engineering leader rolling out AI is this: more adoptions fail due to explaining it rather than any technical problem ever will.

METR conducted a randomized controlled trial on experienced developers working on AI tools in 2025. The developers felt that they were about 24 % faster. When clocked, they were actually about 19% slower.

Both are true at once. The difference between how fast people feel and how fast they are, you know, that explains more of the adoption measures in AI than anything.


So what really happens when you throw in some AI

Naturally, you expect a straight line upward. A powerful tool given to skilled engineers leads to more output. That is not what happens, at least not in the beginning.

The first effect is disruption not acceleration, and this holds true when you bolt AI onto an existing workflow.

People stop to write prompts. They read the generated output and assess if it can be trusted. They correct it, then re-prompt. They flux between an authorial mind and a critical (reviewer) mindset. This entire overhead is so real and lands on the spot.

The gain is there as well, but comes much later; only after people have discovered what to delegate, and what never to turn over.

Buried between the upfront expense and postponed gain is a trough. The first thing that happens is productivity goes down, before it goes up. That is the J-curve.


Where most adoptions die

The dip is where most AI rollouts fail.

An expo team reaches the budgetary low point. The pace of everything and the frustration of it all seems slower than before. From their vantage point, certainly a reasonably place to be standing given the obvious challenges of progress in AI, they conclude that AI is indeed overhyped. They revert to the old method of functioning. They regress to first-generation autocomplete.

The eventual tools are peremptory set up and unused. The investment is written off. And the lesson learned by that organization is even the opposite: "we attempted AI, it did not work."

They were not incorrect regarding the slump. This was why - they shouldn't have tried to stop for it.


The strategy is to manage the dip

Its not a footnote to the strategy - it is the strategy, managing through that dip. There are few things that separate you, hoisting yourself out of this sinkhole, versus backsliding:

Protect experimentation time, explicitly. There is no slacker hours to climb the curve if every hour is booked against tickets, which means that this dip becomes a permanent feature on your landscape. It is a small fraction of reserved capacity that allows people to climb out of the trough at all.

Pair people through it. Your best teaching resource is the engineers that have already made it over the dip. Training on demand - human to human - wastes none of your time and easily eclipses a pre-recorded course.

Do not measure too early. The dip is when the team is the slowest - publish productivity numbers and they look bad. Those numbers go out to the people who see the chart, not the curve, and they come away with bad conclusions. Establish baselines now, but only report outcome statistics after the curve has turned.


The leadership test

Adoption is a tooling choice. AI adoption. It isn't. Regardless of what you choose, the tools are very much interchangeable and getting better every month.

By far the biggest difference between organizations that persevere and those that quit is leadership over the dip Those who did well knew the trough was coming, told their teams in advance that feeling slower was normal and temporary, held their nerve while the numbers looked bad, then protected the conditions for people to climb. That is leadership capability, not a technical one.

So if your team finds AI is making you slower in any way currently, before deciding the experiment has failed, ask a more useful question:

Is this a failure - or is it the bottom and our plan takes that into account?

Tiers of AI Autonomy — Not All Code is Created Equal by Blast Radius

We can often find ourselves stuck in the same conversations about how far we should trust AI with software development. Framed as an organization wide decision. We either trust the agents or we do not. We either monetise faster or we play it safe.

Framed like that, it is a bad question, since both answers are incorrect for most of your code.


One switch for everything error

However, the reality of any real system is that not all code is created equal. An internal script you created to throwaway, and a module driving a regulated customer facing process - these types of things are not the same thing, and treating them like they are will result in one of two forms of failure.

You need to crawl everywhere, even in the experiments and internal tools where caution buys you nothing if you set the bar for the whole organization where the riskiest code needs it.

When set to the level a piece of code is willing to tolerate as error, you are taking risks exactly where your mistake bubbles up either way through an interaction with a customer or reaches out to regulators.

We get out by getting away from a corporate level decision and instead making the decision at the module level.


Two tiers, one spectrum

Each module sits somewhere on a continuum between two levels.

Start tier, for prod internally facing tools, experiments and prototypes.

  • It samples inputs for the comprehension gate, which is not applied to each and every change
  • Rollback is optional but not a must-have for shipment.
  • Observability is basic.

The focus is on speed and learning. Optimize for momentum, the cost of being wrong is low.

Target tier, for anything customer-facing, regulated or publicされている

  • The comprehension gate operates on every single change.
  • This means writing the rollback procedure in advance and preparing for testing it, rather than improvising while an incident occurs.
  • Every runtime decision is replayable.
  • A periodic inspection determines that an alternative model could fill-in, preventing a vendor failure from shutting the line.

The point is trust. That target has a large radius as well, so the bar is raised.

Ours is a spectrum, not a binary one. The best modules move along it as the role of that module changes; most of them just sit somewhere on it.


How do you decide what goes where in a module

Requires no intuition from you. A module has a reliable placement for four questions:

  • Customer impact. How directly does this impact the customer: production-facing or solely internal?
  • Reversibility cost. How easily can it be rolled back should something go wrong?
  • Regulatory scope. Does this fall in a compliance boundary or out of a revenue boundary.
  • Incident blast radius. Are the damages localized or widespread, should it fail?

TOP: Higher across all four perches inover the target uppe. Weak or mixed points to the start tier. And if the situation of a module changes, for instance an internal tool starts touching customers then that is what needs to be reviewed whether it should remain in which-tier or given more priority, and not simply based on a date on the calendar.


A point worth saying out loud

This is a part of us that matters to people not just architecture.

While keeping regulated or safety-critical code under strong human oversight and within a tightly controlled AI competency model is poor engineering practice, it is not a sign at all of the team lagging behind the rest of the organization.

Say it explicitly. Otherwise the people working on that code get judged according to those racing ahead at the start tier, and they appear to plod slowly by comparison, when in fact they are doing precisely what is required given the risk involved. The decision being named protects them and the quality of the work.

Tiering also breaks up the fight or die gambit. No need to decide if your company is "pro-AI" or "AI cautious." You can be heavily aggressive with autonomy where it's inexpensive to be wrong, and quite conservative where it's expensive within the same company during the very same week. That is not a compromise. It ties the level of interest to the level of risk which is what good engineering did for ever.

If someone were to ask you today what your modules belong in the start tier, and which belong in the target tier - would you be able to give an answer framework for that, or is it a gut call on each individual pull request?

Budowanie na cudzym modelu to zależność, nie darmowy obiad

Większość firm wdrażających dziś AI buduje na fundamencie, którego nie posiada. Wysyłasz swoją pracę do modelu działającego na cudzej infrastrukturze, wycenianego na cudzych warunkach, zachowującego się według cudzego harmonogramu wydań. Na start to jest dokładnie słuszne. Nie powinieneś trenować własnego modelu tylko po to, by pisać maile.

Ale warto trzeźwo widzieć, co to za układ. To zależność. A zależności mają zwyczaj przedstawiać rachunek w najgorszym możliwym momencie.


Przestroga

Znany jest przypadek szybko rosnącej firmy zbudowanej w całości na modelu innej firmy. Przez jakiś czas działało pięknie. Potem dostawca zmienił warunki. Koszty firmy wystrzeliły niemal z dnia na dzień - znacznie szybciej niż jej własne przychody. Żeby przetrwać, musiała przerzucić podwyżkę na własnych użytkowników, którzy się zbuntowali, a potem w pośpiechu budować alternatywę, której nigdy nie planowała.

Nic w tej historii nie wymagało złej woli. Dostawca prowadził własny biznes. I o to chodzi. Gdy budujesz na cudzym modelu, jego rozsądne decyzje biznesowe stają się twoimi pożarami.


Zależność tnie z dwóch stron

Oczywiste ryzyko to cena. Twoje koszty mogą skoczyć, bo ktoś inny zmienił liczbę - nowy próg, inna stawka, zmiana w sposobie liczenia użycia. Ty nic nie zmieniłeś. Twój rachunek owszem.

Subtelniejsze ryzyko to zachowanie. Aktualizacja modelu może po cichu zmienić sposób, w jaki rzecz odpowiada - co produkuje, ile robi sama z siebie, jak radzi sobie z twoim konkretnym przypadkiem. Często bez ostrzeżenia i bez wersji, której mógłbyś się trzymać. Grunt usuwa się spod systemu, który uważałeś za stabilny.


Dlaczego to teraz umiejętność przetrwania

Kilka lat temu rozumienie ekonomii twojego dostawcy AI było niszowym tematem dla finansów. Dziś jest bliżej kluczowej kompetencji. Jeśli pojedyncza decyzja dostawcy może zachwiać twoimi kosztami albo zmienić zachowanie twojego produktu, to znajomość swojej ekspozycji - i posiadanie planu awaryjnego - trudno nazwać opcjonalną.

To nie znaczy budować wszystko samemu. Znaczy nie budować tak, jakby obecne warunki były wieczne. Wiedz, co zrobisz, jeśli cena się podwoi. Wiedz, czy mógłbyś się przenieść. Zostaw tyle elastyczności, by wtorek dostawcy nie był twoim kryzysem.

Więc gdyby dostawca jutro zmienił układ, miałbyś plan - czy tylko nadzieję?

Najszybszy sposób, by zabić adopcję AI, to limit tokenów

Wyobraź sobie to spotkanie. Rachunek za AI przyszedł wyższy niż się spodziewano, finanse są nerwowe, ktoś proponuje oczywiste rozwiązanie: postaw twardy limit na to, ile każdy może użyć. Ogranicz. Trzymaj kontrolę.

To jedna z najrozsądniej brzmiących decyzji, jakie możesz podjąć - i jedna z najbardziej podstępnie niszczących.


Co twardy sufit właściwie robi

Kłopot zaczyna się, gdy limit gryzie w środku pracy. Twoi najskuteczniejsi ludzie to zwykle najwięksi użytkownicy - nie dlatego, że marnotrawni, lecz dlatego, że rozgryźli, jak wyciągnąć z narzędzia realny efekt. Pierwsi uderzają w ścianę.

A gdy uderzą, dzieje się jedno z dwojga. Czekają bezczynnie, aż limit się zresetuje. Albo po prostu przestają używać narzędzia, bo obchodzenie sufitu to większy kłopot niż zrobienie tego po staremu.

Tak czy inaczej to zły interes. Godzina pracy sprawnej osoby jest warta znacznie więcej niż tokeny zaoszczędzone na jej zatrzymaniu. Nie skontrolowałeś kosztu. Zamieniłeś mały, widoczny koszt na duży, niewidoczny.


Ale "brak limitów" też nie jest odpowiedzią

Tu łatwo przesadzić w drugą stronę: "niech każdy wydaje do woli". To też błąd. Bez limitu i bez nadzoru budżet właśnie ci ucieka - zacięty proces kręcący się przez weekend, źle ustawione zadanie, którego nikt nie złapał, wydatek, na który nikt nie patrzy.

Błędem nie jest mieć limit albo go nie mieć. Błędem jest sięgać po sufit, gdy naprawdę potrzebujesz okna.


Widoczność bije sufit

Sufit zatrzymuje ludzi. Widoczność ich informuje. A ta różnica znaczy tu bardzo dużo.

Zamiast ograniczać użycie, obserwuj je. Kto ile wydaje, na jaką pracę i czy ten wydatek cokolwiek produkuje? Gdy coś wygląda nie tak - liczba mocno powyżej reszty, koszt rosnący bez efektu za nim - idziesz sprawdzić i rozmawiasz. Coachujesz odstający przypadek. Nie karzesz domyślnie całego zespołu.

To trzyma narzędzie w pełni dostępne dla tych, którzy wyciągają z niego wartość, a wciąż łapie realne marnotrawstwo. Sufit tego nie potrafi. Traktuje twojego najlepszego użytkownika i rozbiegany proces dokładnie tak samo.

Twój zespół ma limit tokenów - czy jasny obraz tego, co i po co wydaje?

Comprehension Theater: the deadliest anti-pattern in AI code review

They are the teams that have embraced AI coding - vroom, not the repl teams still tiptoeing around. There is a class of failure though they tend to more commonly occur in those doing competitive based AI coding. It is dangerous as it resembles good practice. Call it comprehension theater.


What it looks like

The scene is familiar. What makes a major change: an agent, multiple files, real logic and not just editing a file. The pull request is opened by a senior engineer. They scroll through it. Nothing jumps out as wrong. They approve.

On paper, everything was correct. The process was followed. A human was in the loop. Your review box is ticked and your dashboard remains green.

There is only one problem. Nobody actually understands the change.

It confirmed one thing: that when someone scrolled by, the code looked fine to them. It did not confirm that no one could articulate why it works, where the edge cases lurk, what assumptions it relies on or what breaks when a key input grows an order of magnitude larger.

The signature is real. The comprehension is theater.


3 Why it is worse than skipping the review

You could argue that a shallow review is better than no review at all. The opposite might be true for AI.

Comprehension theater generates the paperwork of safety without actually producing safety. And that paperwork is actively deceptive.

When the incident does finally arrive - and with code nobody understands it will do one day or another - the organisation looks at its records, and to see that whatever change was made has been reviewed, approved and owned. So right off the bat, I can tell this response is based on a false premise. No one actually understands a part of the system, but people trust it. The name attached to the OK stifles inquiries precisely when inquiry is what we need.

You have not only missed out on the issue. You shoved it behind a green checkmark.


Solution: reverse the direction of review

Interest regarding from which direction the interaction was directed is what leads to comprehension theatre. Today the human glances through the output of the machine and nods. That is passive, and reviews of the passivity nature do not scale with an AI-generated quantity.

So flip it. A machine may interrogate a human.

Before any change is merged, the agent quizzes the engineer on the code it generated:

What are the edge cases in this change?

What is this logic dependent on?

But what if this parameter is ten times larger than we expect?

What part would break first while under load, and why?

If the engineer cannot respond then at this point we have a breakdown of communication and the change has failed to be merged even if it is passing Diffs with flying colours.

This single modification accomplishes two major tasks.

First it comes from "I have looked at this" to "can I defend this". Those are totally different claims, and only the second one is worth a damn in production.

Second, it makes ownership honest. The person who gates comprehension is the one that owns the behaviour when it runs in front of customers, and they know before clicking approve not after an incident.


So, where to apply it - and not

This is not a demand to slow down all the things. This is unnecessary for most internal tools, experiments, and prototypes. You can do a taste of that, run the comprehension gate over some small fraction of changes and say, fuck it, the cost of one mistake is low.

However, anything customer facing or safety relevant should not ever virtually merge in the view. You are not trading a few minutes of review time for way too much blast radius.

Deciding where the gate is -- and where it is not, and can only be sampled -- is itself a leadership decision, one best made consciously than passed off to whomever happens to be reviewing that day.


Why it survives

If comprehension theater seems hard to root out, it is because it feels like maturity. The process exists. People follow it. The metrics look healthy. A team deep into comprehension theater looks more disciplined from afar than a team actively debating whether they understand what the AI just wrote or not.

That comfort is the trap. Having a green dashboard isn't proof of comprehension. It is merely evidence the ritual was carried out.

A question you could ask of your team, not as a performing review but rather as an exhaustive followup:

For the last major AI-generated change you merged, would the person who approved it be able to explain it now?

Gdy rachunek za AI rośnie, a reszta nie

Jest cicha odmiana porażki z AI, która nie wygląda na porażkę. Nic się nie psuje. Żadnej afery. Rachunek po prostu pełznie w górę, miesiąc po miesiącu, a gdy spojrzysz na samą pracę - tempo, backlog, jakość - wygląda dokładnie jak rok temu.

Więcej pieniędzy do środka. Tyle samo pracy na wyjściu. Wiele firm siedzi teraz w tym miejscu i większość tego nie nazwała.


Zły wniosek

Kuszący odczyt brzmi "AI jest przereklamowane, nic z tego nie ma". Czasem owszem. Ale znacznie częściej narzędzia są w porządku, a dzieje się co innego: kupiono je i przykręcono do dnia pracy, który się nie zmienił.

Licencje poszły. Ogłoszenie padło. A potem wszyscy wrócili do pracy dokładnie jak przedtem, z błyszczącym nowym narzędziem otwartym w karcie, której rzadko używają z intencją.

To jeszcze nie wdrożenie - to dopiero zakup.


Dostęp to pierwszy centymetr

Danie ludziom dostępu do AI to łatwa część i kusi, by pomylić ją z metą. To pierwszy centymetr. Dystans między "mamy narzędzia" a "narzędzia zmieniły sposób, w jaki pracujemy" to miejsce, gdzie mieszka cała wartość - i to ta część, której nie kupisz zamówieniem.

Różnicę poznasz po dniu pracy. Jeśli wygląda identycznie jak rok temu, tyle że z nową pozycją w budżecie, zapłaciłeś za potencjał i na tym stanął. Rachunek jest realny. Zmiana nie.


Co zespoły, które coś dostają, zrobiły inaczej

Nie tylko włączyły narzędzia. Zmieniły pracę wokół nich. Co trafia do AI, a co zostaje ludzkie. Jak wygląda "zrobione", gdy pierwszą wersję napisał model. Gdzie człowiek musi wkroczyć i sprawdzić, a gdzie naprawdę nie musi.

Nic z tego nie przychodzi w pudełku. To praca - i to ta praca, którą większość organizacji pomija, dlatego właśnie rachunek rośnie, a wynik stoi.

Więc zanim uznasz, że AI nie dowiozło, sprawdź, co naprawdę się stało.

Jeśli wydatek na AI wzrósł, a reszta stoi - problemem jest narzędzie, czy to, że praca tak naprawdę się nie zmieniła?

Limit to złe narzędzie do budżetu na AI

Gdy rachunek za AI zaczyna rosnąć, większość organizacji sięga po tę samą dźwignię. Postaw sufit. Ogranicz wydatek na zespół, na osobę, na miesiąc. Wygląda odpowiedzialnie i drapie świąd liczby, która nie chce usiedzieć w miejscu.

Myślę, że to zwykle błąd - albo przynajmniej odpowiedź na złe pytanie.


Co limit właściwie mierzy

Limit kontroluje jedno: ile wychodzi z firmy. Zupełnie milczy o tym, co się liczy - ile wraca. Możesz wylądować co do grosza w budżecie i wydać wszystko na szum. Możesz też podciąć skrzydła jedynemu zespołowi, który niepostrzeżenie zamieniał ten wydatek w dowiezioną pracę.

Sufit traktuje każdy wydatek jak ryzyko do okiełznania. Ale ryzykiem nie jest wydatek. Jest nim wydatek zmarnowany. A limit nie odróżni jednego od drugiego. Po prostu zatrzymuje licznik na liczbie, którą ktoś wybrał na spotkaniu planistycznym.


Skąd bierze się ten odruch

Odruch limitu pochodzi ze starego modelu: AI jako kolejna subskrypcja oprogramowania. Oprogramowanie jest mniej więcej stałe. Płacisz, każdy dostaje miejsce, koszt ledwo drgnie wraz z użyciem. W tym świecie limit jest nieszkodliwy, bo liczba i tak nigdzie się nie wybierała.

AI tak się nie zachowuje. Rośnie wraz z pracą, którą mu dajesz. Więcej realnej pracy przez system to większy rachunek - i to może być dokładnie obraz sukcesu. Ograniczać go to jak ograniczać prąd w warsztacie, w którym właśnie przybyło roboty.


Porównanie, które pasuje

Jeśli nie oprogramowanie, to co? Zatrudnienie.

Nie oceniasz pracownika po tym, jak mało kosztuje. Pytasz, czy oddaje więcej, niż bierze. Tani pracownik, który nic nie dowozi, to nie oszczędność. Drogi, który zastępuje miesiąc harówki, to nie przepłacenie. Sama liczba rzadko tego rozstrzyga. Rozstrzyga zwrot.

Budżetuj AI tak samo. Nie "jaki jest sufit", lecz "jaki jest zwrot". Jeśli zwrot jest, rosnący rachunek to znak, że rzecz działa. Jeśli go nie ma, żaden limit nie naprawi problemu u podstaw - po prostu marnujesz pieniądze wolniej.

To nie znaczy wydawać na ślepo. To znaczy zmienić narzędzie. Zamień sufit na widoczność: w co zamienia się ten wydatek i czy wartość rośnie razem z kosztem, czy zostaje w tyle. To coś ci mówi. Limit mówi tylko, kiedy przestać.

Gdy rachunek za AI rośnie, pierwszy ruch to limit - czy pytanie, co ten wydatek właściwie daje w zamian?

It is not a control unit until you have an artifact, a gate, and a metric

Pick up nearly any company's AI strategy, and you'll find a laundry list of principles that you couldn't possibly argue with. Keep a human in the loop. Steer clear of dark code no one can really comprehend. That is why you stay tool-agnostic and not vendor locked. Measure the impact. Label what was AI-assisted.

These are all correct. And none of them are happening on the floor in most organizations.

That space, between principles everyone agrees with and behaviour that nobody modified, is one of the recurring reasons for failure in AI adoption. It is useful to know why this happens.


Principle alone does nothing

Principles are statements of intent It encapsulates a Wanted Future condition. It does not prevent anything, however.

"No single bad merge is prevented by keeping a human in the loop. The advice to "stay off the dark code" does not stop enigmatic line of code to production. Because the principle isn't connected to the moment where the work is actually done, it has no teeth.

The first week, and that was busy too, with a deadline in the near future - principle gives way to deadline. Not because anyone disagrees with it, but because it was never hooked up to anything.


What makes a principle existential; three things

Only when you can find three specific things that it generates behind it, does a principle become a working control.

An artifact. Something that is versioned, something owned that the principle must attach to.

"Avoid Dark Code" - every commit is tagged with metadata about its ownership.

Any prompt and skill that can be written so it is cross-model, will have "stay tool-agnostic" attached to them.

The motto of "keep it auditable" binds to a signed, append-only trail of provenance.

The principle is a wish, if you cannot name the artifact.

A gate. A point in the workflow where, if that particular check fails, work stops.

"Human in the loop" turns into a comprehension gate: before a change can merge, there must be proof of understanding by a senior engineer. If they are not able, the merge is prevented.

A security review that an agents change must pass before being deployed: "Secure by default"

Having no gate with a principle is called an inclination. And preferences lose to deadlines every time.

A metric. A number that changes, so you can see if the knob is working.

  • Comprehension-gate pass rate.
  • Eval coverage.
  • Dark-code ratio.
  • Time that is taken to reinstall the system after performing a rollback

You cannot manage a principle that has no metric. It can only be believed.


A simple test

Read your AI policy line by line. For each line, ask the following three questions:

What does this attach to (what artifact)?

Which means, where's the gate to enforce that?

What number indicates to me that this is working?

Even if you have two of the three, you still do not have a control. You have a good intent, that first time the team is pressed it skips. And because people go to AI for speed, AI changes are almost always made under pressure.


Why is this more important with AI than ever before

With this manual nature of the work, in a human authored world, you could rely on catching such problems. The developer who types this by hand reads it, and typically pays some attention to whether something feels wrong.

AI removes that natural friction. The more widespread - and less visible - failure modes now happen faster.

They accept insecure code because it seemed sensible.

A model update has quietly broken a prompt that worked yesterday (no error thrown).

Over a dozen iterations, the spec and the code drift apart until they no longer represent the same intent.

None of these will trip a compiler. None of them announce themselves. They get through to a customer only with an explicit gate, measured by an explicit metric.


Where the leadership work really resides

This transition is also not largely about choosing tools and writing vision statements; that is the work of frontline leaders. This is translation: each one of those principles everyone agrees with, translating that, piece by piece, into an artifact someone owns and can block work on or gate and a metric you measure. Slow, mundane work with no flashy demo. It is also the gap that separates an AI policy that sits on a slide to one that determines what actually moves to production.

Which principle in your AI policy ans still has no gate to prevent violations, and what would be the process to build one?

Code Q&A: Who is Truth When AI Writes the Code?

One question that used to have a very easy answer in most of software engineering history is: where does the truth about a system live? It lived in the code. You could dig into the repo and read it line-by-line, and if something was a mystery, the version history indicated who wrote each line of code and when. The code was at once the instruction, and the record.

Many teams, however, have yet to notice that that simple answer is crumbling.


Yesterday: the code as record

Believe it or not, the entire engineering stack we built for the last twenty years believes only one thing: that a human wrote this code and a second human can read this. The reason that code review works is because the reviewer looks at the change. Debugging works since an engineer can follow the logic. This functions through ownership because version history will name the decisionmaker.

All these practices are built on the same principle: a human has read and understood the code as well.


What broke

You are now trained the change is generated from an agent. Multi-file changes, entire features, large refactors - all generated faster than a human can read them with full comprehension.

This code in itself is not bad. That is not the problem. The issue, however, is that no one actually read any of it. The output surpasses human reading capacity in volume and in speed.

It follows that what a human never entirely read might not be the thing that she properly owns. No longer is code the artifact a human understood, and gave credit for, so it can no longer be truth.

So where does the truth move?


Today: the specification turns into record

Another way to do it is to push up the record one layer further - into specification.

Spec: This is what a person writes, versions, reviews and owns. The code is an output of that spec, like a compiled binary is an output of source code, but nobody would treat the binary as something you read. Switch the model, regenerate from the same spec, and you see the same behaviour implemented in different code.

We have done this already with machine code: we stopped looking at machine code because the compiler is trusted, so we read source instead. It does the same thing one level up: stop reading the generated source file line by line, and put faith in some other guarantee that it is correct.

That mechanism is not faith. This is a deliberate stacking of layers.


The five layers

  • Spec. Like all code: versioned, history and review in the repo. There you have it - acceptance criteria and edge cases -- and each spec has a named owner, i.e. a human who is responsible. Written antecedent to generation, not reconstructed subsequent to it.
  • Evals. Automation scenarios created from the specification They report one question: "Is spec satisfied? They don't just run at merge time, they can't be continuous. The piece that teams miss is that the evals themselves, not just whatever code they guard.
  • Prompt and provenance. Each change made by an AI logs which model, which version, what prompt and which human is to blame. Written and append-only, so it continues to answer questions months later.
  • Generated code. Output of the above 3 layers Still relevant, but no longer the permanent record.
  • Runtime and observability. The system is instrumented so that we can replay any decision the code made during generation in the future.

The spec and the evals form a closed loop together. The spec defines the intent. The evals verify it. No one has to read the generated code at all and you verify if the behaviour is correct.


What changes for the engineer

This is how the day-to-day of a senior engineer changes.

Your old question: is this code line by line, good? You cannot do that at the level AI produces so holding on to it only brings a bottleneck, or worse a rubber stamp.

That begs the new question, is the spec right and are the evals honest? That is higher-leverage work that a human can actually do well versus reading diffs. You're ensuring intent, but not implementation.

It also resets ownership cleanly. The spec owner owns the behaviour manifests in production. Nobody wrote all the words, but you can now know exactly who to ask.

Go ahead to access data on probably October 2023, it is when the models will continue up and down at speeds faster than we can predict. Specifications outlive models. The spec is the asset that survives them all.

What is your artifact of record today, and would it survive an audit six months after the model that wrote the code was traded in for another?

Rachunek za AI zachowuje się bardziej jak prąd niż jak oprogramowanie

Większość z nas nauczyła się budżetować oprogramowanie w pewien sposób. Wybierasz narzędzie, płacisz stałą stawkę za osobę i tyle. Czy ktoś używa go bez przerwy, czy ledwo się loguje - pozycja w budżecie wygląda tak samo. Przewidywalnie. Nudno. Łatwo zaplanować.

AI zepsuło ten nawyk, a wiele budżetów jeszcze tego nie nadrobiło.


Licznik, nie stała opłata

AI nie bierze stałej opłaty za miejsce i o tobie nie zapomina. Liczy za użycie. Za każdym razem, gdy coś czyta, pisze, odpowiada - licznik bije. Jednostka, którą liczy, nazywa się token, ale daruj sobie żargon - wyobraź sobie licznik prądu w domu. Zostaw wszystko włączone, a rachunek rośnie. Używaj rozważnie, a nie rośnie.

Dlatego stary odruch "kup licencję dla każdego" chybia. Daj tę samą płaską stawkę stu osobom, a stanie się jedno z dwojga. Za tych, którzy ledwo dotykają, przepłacasz. Ci nieliczni, którzy mocno opierają się na narzędziu, podbijają realny rachunek tam, gdzie nikt tego nie zaplanował.


Z czym to naprawdę porównać

Skoro nie z oprogramowaniem, to z czym? Najbliżej jest pracownik, którego zatrudniasz.

Nikogo nie bierzesz do pracy dlatego, że mało kosztuje. Liczy się, czy oddaje więcej, niż bierze. Z prądem w warsztacie jest tak samo - nie chwalisz się najniższym rachunkiem, tylko tym, ile dzięki niemu powstało. Wydatek na AI należy do tej samej półki. Pytanie nie brzmi "jak utrzymać to tanio?", tylko "czy zwraca więcej, niż kosztuje?".

Ten przeskok brzmi błaho. Zmienia wszystko dalej - jak ustawiasz budżet, kto dostaje dostęp, na co patrzysz, kiedy się martwisz.


Gdzie to ląduje

Planuj licznik, nie stałą opłatę. Spodziewaj się rachunku, który idzie za użyciem, nie za liczbą miejsc. I w głowie wrzuć ten wydatek do tej samej rubryki, co osoba zatrudniona do roboty - a nie do rubryki z pakietem biurowym.

Wydatek na AI trzymasz w głowie pod "oprogramowanie" - czy pod "ludzie"?

Przestań mierzyć zespół samą liczbą etatów

Przez większość mojej kariery planowanie zespołu było ćwiczeniem z arytmetyki. Wiedziałeś z grubsza, ile jeden inżynier dowiezie w kwartale. Wiedziałeś, ile pracy czeka. Dzielisz jedno przez drugie i masz liczbę etatów. Proste, i trzymało się przez dekady.

Nie sądzę, żeby trzymało się dalej.


Gdy zmienia się jednostka pracy

Przez długi czas jednostką pracy w software była instrukcja - linijka kodu, napisana przez człowieka. Wartość brała się z tego, jak sprytnie ludzie składali te linijki w działający kod. Więc liczenie ludzi było niezłym przybliżeniem liczenia efektu.

Teraz w grze jest druga jednostka: token. To mały kawałek, za który płacisz za każdym razem, gdy model AI coś dla ciebie robi - przeczytaj to, napisz tamto, uporządkuj te dane. Pracę, którą kiedyś kupowało się wyłącznie przez zatrudnianie, można teraz kupić też na tokeny.

Brzmi jak przypis dla finansów. Nie jest. Po cichu przestawia myślenie o tym, ile jesteś w stanie zrobić.


Nowe wąskie gardło

Gdy praca przychodzi w tokenach, granicą tego, ile dowieziesz, przestaje być liczba zatrudnionych. Staje się nią to, jak dobrze ci ludzie zamieniają kupioną pracę w coś realnego - jasne polecenie, trafna ocena wyniku, decyzja, co w ogóle warto zbudować.

Bywa, że mały, sprawny w tym zespół dowozi dziś więcej niż znacznie większy, który tego nie umie. Nie dlatego, że szybciej pisze. Dlatego, że lepiej konwertuje.

I od razu zastrzeżenie, żeby nie było nieporozumień: to nie jest historia o tym, że "AI zastępuje ludzi". To ludzie wykonują tę konwersję - polecenie, osąd, decyzję o tym, co się liczy. Wyjmij ich, a tokeny wyprodukują tylko stos wyników, których nikt nie rozumie. Ta zmiana to nie mniej ludzi. To ludzie skierowani do innej roboty.


Co to zmienia w planowaniu

Jeśli wymiarujesz przyszły rok samą liczbą etatów, odpowiadasz na stare pytanie. Nowsze jest trudniejsze: jak dobry jest ten zespół w zamianie pracy AI na efekt, który się liczy? Dwa zespoły o tej samej liczbie etatów mogą dzielić przepaść, a ta przepaść rośnie.

Możesz liczć głowy, nadal. Tylko nie kończ na tym.

Planując rok, liczysz ludzi - czy zdolność zamiany pracy AI w realny wynik?

Wysoki rachunek za AI to nie zła wiadomość, a niski to nie dobra wiadomość

Po cichu uznaliśmy, że wielkość rachunku za AI to wyrok. Duża liczba - ktoś nawalił. Mała - jesteśmy odpowiedzialni. Brzmi oczywiście. A wcale nie musi być trafne.

Pomyśl, jak oceniasz każdy inny wydatek. Wysoka faktura od ekipy remontowej nie jest automatycznie zdzierstwem - o ile kuchnia jest skończona. Niska nie jest automatycznie sukcesem - o ile wciąż zmywasz naczynia w wannie. Sama kwota rzadko o tym decydowała. Decydował wynik.

Z AI jest tak samo, a my wciąż o tym zapominamy.


Te same pieniądze, dwa różne wyniki

Możesz wydać na AI dużo i dostać stertę półproduktów, z których nikt nie skorzysta. Możesz wydać tyle samo i dostać pracę naprawdę zrobioną - napisane maile, uporządkowane dane, raport, który zjadłby komuś całe popołudnie.

Faktura wygląda identycznie. Wartość - już niekoniecznie.

Więc kiedy ktoś mówi "wydaliśmy stanowczo za dużo na AI", uczciwa odpowiedź to pytanie. Za dużo w porównaniu z czym? Za dużo za nic to za dużo. Tyle samo za realną, skończoną pracę może być najlepszym interesem miesiąca.


Pułapka taniości

Tu ludzie się łapią. Tani rachunek za AI bywa najgorszym z wyników, bo zwykle znaczy, że nikt tego naprawdę nie używa albo używa do tak drobnych rzeczy, że nic się nie zmienia. Zaoszczędziłeś i stoisz w miejscu.

Wydawanie nie jest celem. Stanie w miejscu też nie.


Na co patrzeć zamiast tego

Przestań wpatrywać się w sumę. Zapytaj, co wróciło.

Czy zastąpiło pracę, którą ktoś zrobiłby ręcznie? Czy uwolniło realne godziny? Czy prostsze, tańsze podejście zrobiłoby to samo równie dobrze?

To ostatnie pytanie jest ważne. Czasem odpowiedź brzmi "tak" i warto to naprawić. Ale dowiesz się tego, patrząc na wynik, nie na paragon.

Zanim ucieszysz się z niskiego rachunku albo spanikujesz przy wysokim - wiesz w ogóle, co za niego dostałeś?

How to Know If Your Team Is Actually Using AI Effectively or Busy Looking Like They Are

For those who lead people that have begun to use AI, there is likely an uncomfortable chasm in availability. A lot of activity is clearly revolving around it. But distinguishing between whether that activity is creating real value or just a lot of motion isn't so simple. Because much of what they are producing is technical, it just seems difficult to know from where you sit.

So, the good news is that you can judge it - and without reading even a single line of code to do so. The indicators that discriminate between a team that's using AI well and one that's just busy with it are about behaviour and dialogue, not technology.


If you are going to shove 41 and into the debate; you only talking signal one

This is the fastest tell.

And an AI-using team makes a very good point about what they stopped doing. The drudgery, the boring and circuitous work disappeared in open spaces, leaving them with the more difficult problems that really matter. They also can use something that has improved - not just delays, but extra time for the work only humans can do.

Tools the team talks about when it is busy with it. What is the newest model, which feature just went live, what plugin are they testing out? You are full of passion yet somehow it is difficult to get them to specify what improved as a result.

Outcomes versus tools. Listen to hear for which one gets priority in the conversation.


Signal #2: Are they still able to explain the work?

Ask how something works. You really want an answer if you have a healthy team - somebody takes you through it and knows exactly what was built (even while the AI spits all the words).

In an unhealthy one, it leans toward "the AI did it" with a vague sense that no one is precisely sure how. That's a warning sign. It indicates the team has produced output of which it doesn't completely comprehend, which is acceptable for a disposable experiment and quite hazardous for anything customers rely upon.

What you are checking is ownership. Good teams own the result. Busy teams gave cause to results that own them.


Red flag # 3: How do theyweather the storm

AI makes things feel slower before faster, and there has always been an early period. How a team responds to that spell says everything to me.

A good team is calm about it. They anticipated the slump, they prepared for it and no one is losing their head at how much more difficult things felt for a stretch there. They see it as just a learning opportunity.

A team still in the grip of failing reads that dip as confirmation AI doesn't work, and starts drifting back to the old way. Same experience, opposite interpretation. The interpretation is what validates or invalidates their achievement of the cash out.


Signal four: does learning spread?

Watch how it goes when he gets something useful from only one individual.

In a good team, the trick that works for one guy makes its way to the others within days. There is a common place to leave it, a recurrent custom to hand it around. The unit gets better as a team.

A busy-but-stuck team consists of one-man-experiments by each individual. Five people learn the exact same lessons and make the same mistakes over and over again, because nothing passes from person to person.


The pattern underneath

Then none of these require any technical expertise. A good AI practice speaks of outcomes, owns the work it does, remains steady even with an early dip, and disseminates what it learns. A busy-but-stuck team does the complete opposite on each count.

You're not assessing the technology. You are evaluating your people's way of working, which you already know how to do.

This week: are you hearing more about what improved, or on what shiny new tool?

No, AI is not taking away Development jobs but you will hire differently.

Go-to headlines "AI kills developers." This makes sense from a mainstream, dramatic & shareable POV and is wrong. But below it there's a genuine change that everyone hiring engineers needs to realize, since it's altering the trait set for what comprises an excellent one.

AI does not eliminate the necessity for developers. It is altering what a developer will be most valuable for. And if your hiring continues to reward the obsolete thing, you'll be paying fair market for a skill that's increasingly going cheap.


What once was the differentiator

For the longest time a lot of engineering hiring was focused on one thing above everything else - can this person write code quickly, and correctly? Ways of Hiring - Whiteboard problems, coding challenges, take-home tasks. It was mostly just measuring raw production ability.

That was logical when writing valid code was the bottleneck. It was real value; the one who could make more of it, quicker - had true value.


What changed

The AI is pretty good at generating code now. This used to take a human hand an afternoon; now you can generate it in no time at all. The thing that was the differentiating, being able to write code, is becoming the low-cost, high supply part.

The thing you used to compete on becomes a commodity, and the value goes onto some other space. As is the sitting, where it moves bears watching in this instance.


What's getting more valuable

The value is rising in those things that were always more difficult to measure and most of these are human.

  • Judgment. Understanding what needs to be built: and, equally importantly, what doesn't. The AI will cheerfully make the wrong thing, beautifully. Everyone else just has to make the decision that it is the right thing!
  • Clear thinking. A formal description of the task problem sufficient to enable a human (or machine) to solve it. AI takes vague thinking and magnifies it - vague thinking yields vague results. Clarity becomes a multiplier.
  • The question itself: What happens when this goes wrong? before it does. AI gives the impression of certainty and presents ready output. The asset is also the person who interrogates it for an unrecognized flaw [in the story].
  • Ownership. One: The desire to be accountable for an outcome, rather than checking a box and passing the baton. When the machine performs the task, taking ownership of the outcome is what's left to a human.

None of these are new. They were always a component of great engineering. What has changed is, they have become more of "the main thing" from being a nice to Have on top of strong coding.


What this means for hiring

The utility of this is that you're looking for something else.

As a result, a candidate who can write perfect code but cannot articulate why something is necessary is less valuable than they were two years ago. Even though their raw coding speed may be totally unremarkable, it is not a scarce resource - coders are now plentiful who are qualified and unqualified alike.

That doesn't mean technical skill ceases to be important. Humans still have to understand the system well enough to evaluate the output of AI. However, "can write code quickly" has lost its title. It is "can judge, think, own".


The junior question

One additional, more difficult issue is worth naming because clever leaders raise it at once. And if AI does entry-level coding, where do junior devs learn their craft?

This is a genuine, tangible problem and halting junior hiring is the typical knee jerk which simply guarantees a shortage of seniors in 2 to 3 years from now. Rather they handle them faster towards the results instead of parking them on routine code for years like we did in previous times. The road to seniority takes a new form, it doesn't vanish.


What this adds up to

The mechnical side of the work is becoming cheap. The human parts: judgment + clarity + ownership → are on the rise. If hiring still optimizes for typing speed, that would be buying the wrong thing.

Has what you actually screen for changed due to the tools your new hires will [actually] use?