What am I Paid for?

What am I Paid for?

Now, if you ask me in my personal time about the thing we are calling AI, Generative artificial intelligence (GenAI), I would list off one of the many, many reasons that I refuse to use the giant end-game-capitalism copyright-theft machine.

However, I am employed in a technical capacity, and that means during my job I have to leverage GenAI in order to stay relevant.

I don’t tend to talk about it publicly, because there are so, so, so many posts about it already. You are not short of other things to read about how it’s going to reshape the human condition, or murder your grandmother with a Greggs Sausage Roll at the first opportunity.

However, I do want to talk about my recent usage at work.

Reframing what am I paid to do

One of my first reactions to seeing tools like Copilot generate code, was similar to many other developers; that my job was in danger. If the GenAI tools can write an entire app in under an hour with a sufficiently good prompt, what will I do until I retire? I need to pay for a roof over my head. Do I need to retrain as a plumber?

But the more I have spent time using GenAI recently, the more I wondered if my job was ever to just write code.

Sure, I spent a decade where the primary output of my day was code, but to know what to write I had to do requirements gathering, and made sure that what I built met those requirements. I spent huge volumes of my time writing tests, to ensure I could detect regressions in behaviour.

I think I was actually paid for behaviour.

Specifically, I was accountable for building something that exhibited the behaviour the business wanted; something that met the requirements.

Now that the giant water and energy vacuum that is GenAI can write the code, it might change how I produce and assert that behaviour, but not that I am still accountable for it.

Code was how I made something that exhibited the behaviour demanded by the requirements.

But is it still?

Accountability

I heard a fascinating story somewhere earlier this year1.

You are invited to an important meeting where an important decision is set to be made.

This decision, if executed will make the business a lot of money. If not executed correctly, the business will lose a lot of money.

Unfortunately you are busy, so you delegate attendance and note-taking to your brother, who is well known to make mistakes.

He attends, but his handwriting is horrible, and you miss something important. Your business loses a lot of money.

Who is accountable?

Now, this is improbable, because your brother would spend most of his time trying to install Teams on his laptop before missing the meeting because a plugin was not installed correctly.

But GenAI is known to make mistakes — a warning put underneath every chat box — so if you delegated anything to it, do you think you could pass the blame to it?

Even with GenAI, the person using it still holds the accountability, so developers using it to write code, are still accountable for the behaviour of the code. No self-respecting business would accept you blaming the AI because it didn’t do what you wanted it to.

I think if businesses are OK not blaming their employees when data breaches happen because GenAI decided to do authentication client-side (as it was not just more efficient, but logically simpler), then we don’t need to ask ourselves the question:

How do we, as developers, leverage GenAI to accelerate delivery (which is demanded by almost all boards of directors2), but also give us the certainty that the often error-filled output of GenAI does what we want it to do, so we are happy being held accountable?

Behavior-driven development

Behavior-driven development (BDD) is something which I have used most of my career, and I think it’s more important than ever, now that our connection with the ‘production’ code is weakened.

If we look at GenAI as a new starter/a graduate, we don’t want the code they commit to fly straight into production. It is very important that a second pair of more senior eyes goes over what they output.

But what if this new starter is generating code faster than you can review? They type so fast, reworking entire sections of the codebase can be done before you get back from the bathroom?

What protections could you put in place to ensure that you can benefit from this speed (because your backlog is always looming), but also stay accountable for the output of this new starter?

I think — and I am still working through this — that BDD is the answer.

We let this immensely fast, never-tiring, new starter control all the ‘production’ code, and all the unit tests (because they are tightly coupled with the production code), but we make them also output human-readable BDD test scenarios, which we can validate faster3:

Given I'm an author user  
And I have an article  
When I go to the article page  
And I press the publish button  
Then I should not see the error message  
And the article should be published

The code that executes these steps (Fixtures), tend to be quite simple (assuming you are writing single use BDD steps):

from pytest_bdd import scenario, given, when, then

@given("I'm an author user")
def author_user(auth, author):
    auth['user'] = author.user

In my experience, forcing GenAI to write its integration tests in this way, allows me to ensure the code has the intended behaviour. I review the fixtures, and the scenarios, and I completely ignore the production code and unit tests.

I admit — as an incredible engineer I respect pushed back on this week — that there is a greater degree of risk here. His main point was that we wouldn’t test the finer minutia of behaviour, which I agreed with. Who would test background colour?

Up until now, there was a broad range of ‘common sense’ behaviour which we just assumed a developer would write well, because of their training.

But how do we train engineers — whose university did not teach your code style preferences — when they join?

We would spot them using System.out instead of log.info in a code review, and we would correct them. Then, when it keeps happening every time you hire someone, you would automate that correction, so you didn’t need to keep it in your mind when doing code reviews.

They would write a if statement slightly differently than you, but as long as it wasn’t overtly overcomplicated, and it did the right thing, you would probably wave it through.

Small deviations were always fine, as long as the behaviour was correct.

I think, now that GenAI can output whole applications in a day, we need to assert the behaviour of the finer minutia of applications, which is even easier now those fixtures can be generated in seconds.

Spec-Driven Development

At work, I have been using Openspec, one of many Specification-grounded, Spec-Driven development (SDD) GenAI agent frameworks.

In short, someone writes GenAI agentic prompts which direct GenAI to work through a typical engineering flow, and to store records in a given format, so future work can pick up where you left off.

I particularly like Openspec, as the chosen record format is very similar to a BDD file; lots of scenarios (Given, When, Then). This means you, as a person, can read through them to ensure they capture your intent before you spend $300 in credits building the wrong thing.

Quite a few of the SDD frameworks store specifications in long, winding, incomprehensible markdown files, which makes understanding expected behaviour hard.

Having the specification written as scenarios means you can check the feature files and fixtures to ensure you have tested everything.

Is it perfect? No.

But it’s a lot better than just having the GenAI forget all the previous behaviour you built. It’s particularly useful as agent will pick up inconsistencies during build.

Enjoyment

I think it’s safe to say that this change in my thinking is a form of coping mechanism.

Generative AI is sweeping through my industry and reshaping job roles at every level. I didn’t complain when I wrote software to automate the work of others, so I think I lose the moral high ground when someone does it to me.

It still sucks, however.

I enjoy coding, and I don’t use ‘AI’ at home because I want to feel a sense of achievement when I do something; whether that’s writing this blogs CSS, or writing fiction.

Doing something with ‘AI’ would be faster, but I don’t feel achievement when I get it done. Many people still do, and that’s great; I’m happy for you.

I used to call writing software an art form. Making the computers complex systems come together to do something specific, akin to a conductor.

But the march of technical progress has swallowed my chosen profession whole, leaving me wondering if development will soon be forgotten like many other ‘heritage crafts’.

Shadow Behaviour

One final thought on my entire post:

Developers not reading and understanding the code obviously comes with the issue of shadow behaviour; features added into to the code which are additions by the GenAI agent, which it does not care to mention, or test. For example, a bypass of your login page to make automated validation easier on /only-llms-please. If you don’t look at the code with your eyes, would you ever find it?

I don’t know if there is any way to pick this up beyond an adversarial GenAI review (and lord help us if they both agree).

Whatever business is driving for this magnitude shift in delivery pace, has to be fine with the additional risk of it.

They probably will still blame the developers.


  1. If you know where this is from, let me know. ↩︎

  2. There is another blog I wanted to write about top-down and bottom-up innovation. ↩︎

  3. Example from the pytest docs. ↩︎