Story 005 - I Tried to Give My Job to a Machine
Published:
There are engineering problems you discover because they are difficult.
And then there are the ones you discover when you are really, really lazy.
For me, it was QA.
I was temporarily assigned GUI testing work to learn a large enterprise application.
The assignment made sense. The business application had complicated workflows, hidden dependencies, and enough business rules to punish anyone who clicked through it without understanding the correct process. Skipping even one step could trigger a painful list of validations later on that would make anyone reconsider their employment.
There was only one problem.
I did not want to keep doing this testing, especially not manually.
The Job
Each testing request followed roughly the same routine.
A developer or product owner would explain what had changed and what needed to be tested. To test the change, I had to reconstruct the existing business process, reach the stage where the new functionality had been introduced, test multiple scenarios, capture screenshots, and prepare a .docx evidence file. This was the only part of my job I hated. The clicks were repetitive, but the thinking around them was not.
Fortunately, I had already built a RAG chatbot (the one I built in Story 000) using the application’s user manuals, FAQs and training video transcripts. Whenever I did not know how to complete a business flow, I asked the chatbot, and it told me what to do and how to do it.
But I still had to perform the manual labor. And again, I hated how much of my time and effort disappeared into something that was pretty meaningless to me and my future.
My Completely Reasonable Response
I decided to build my replacement, a machine that would do this job for me.
The first version of the idea was simple:
Give a Vision Language Model capable of Computer-Use the current screenshot
-> Tell it what needs to be achieved
-> Let it decide the next action
-> Orchestrate the action
-> Capture the new screen
-> Repeat.
OpenAI’s GPT-5.4 and its advertised computer-use and vision capabilities were attractive, but I needed a working proof of concept before management would approve a paid API subscription. So I went looking for an open-weight VLM with computer-use capabilities that could fit inside 12 GB of VRAM.
That search led me to Qwen3-VL-8B-Instruct. I downloaded the model, loaded it through LM Studio, and started with an isolated validation. I gave the model screenshots of the application and asked questions such as:
- What atomic action is required next to complete this goal?
- Has the previous action changed the interface as expected?
The answers were good enough to prove that the model understood the screen and the goal. So I gave it a precise system prompt and forced it to return exactly one atomic action at a time.
I had already learned that asking a model of this size for an entire sequence of actions was a fantastic way to let one bad assumption ruin everything that followed. One action at a time at least limited the damage to one click.
Version 1: A Brain Without Motor Control
The VLM could usually identify what needed to happen next and which GUI element should be used.
Then I asked it for the exact coordinates. It failed badly.
The model might correctly decide that a particular button had to be clicked or a particular field needed to be filled, but then provide coordinates for a nearby label, a surrounding container, or something completely unrelated. I had the execution brain, but I was missing the nervous system and motor control.
Version 2: Giving My Replacement Eyes
Another rabbit hole led me to Salesforce GPA-GUI-Detector, a YOLO-based model for detecting interactive GUI elements.
Perfect.
This would become the eyes of my replacement. The detector returned the bounding boxes and coordinates of visible elements, but it did not answer what text appeared inside those boxes.
For that, I added TrOCR.
Now the perception service could detect a GUI element, crop its bounding box, read the visible text, calculate its centre coordinates, and represent it inside a structured UI map.
If no text was detected, the element received the highly informative label: icon. Not perfect, but good enough for the prototype. The demonstrated flow mainly involved labelled buttons and input fields rather than unidentified icons.
The WSL perception-service code converted each screenshot into a map containing:
{
"id": 17,
"text": "Save and Next",
"center": [1450.5, 812.0],
"box": [1380.0, 785.0, 1521.0, 839.0]
}
Naturally, I then tried sending the screenshot, the complete UI map, the system prompt, and the active goal to the VLM in one request.
That was expensive on my hardware.
Even after paying that cost, the model continued hallucinating coordinates. Lower temperature and controlled sampling did not magically turn the VLM into a geometry engine.
The huge UI map was probably not helping either. That gave birth to version three.
Version 3: Stop Asking AI to Do Geometry
After a vibe feasibility check, I restricted the VLM to selecting only:
- the next mouse or keyboard operation;
- the visible element on which the operation should be performed; and
- the value to type, when required.
A typical response looked like this:
{
"intent": "click",
"target_text": "Save and Next",
"type_value": "",
"reason": "The required fields have been completed."
}
I removed coordinate resolution from the VLM entirely. Instead, I built a deterministic actor that consumed two things:
- VLM output: Semantic action and target text
- UI map: Detected text, bounding boxes, and centre coordinates
The actor attempted an exact text match first. TrOCR occasionally made spelling mistakes, so when no exact match existed, the actor used fuzzy matching.
Then I discovered that the application contained multiple search bars, repeated labels, nested controls, and several elements with identical text.
My first geometric rule preferred the larger bounding box. That rule confidently selected broad containers instead of the actual control.
So I reversed it. The actor began preferring smaller actionable regions and used basic 2D coordinate geometry to resolve remaining ambiguity.
It was not exactly revolutionary mathematics. It was rectangles, text similarity, and stubbornness.
But it worked.
The VLM became the brain. It decided what action should happen and which visible element should receive it.
The perception service became the eyes. It mapped the screen onto a 2D plane containing detected elements and coordinates.
The deterministic resolver became the motor-control layer. It translated semantic intent into a physical target.
PyAutoGUI provided the actual mouse and keyboard actuation.
And then my replacement started using the application.
Unfortunately, Clicking Was Only Half the Job
While the execution subsystem itself was fully autonomous once given a task, the overall system workflow remained semi-autonomous. Give the execution engine an ordered set of business goals, and it could observe the interface, choose actions, resolve targets, and interact with the application entirely on its own.
But I still had to read the request, understand the new change, reconstruct the existing process, and prepare those goals manually.
Apparently, replacing myself required automating the thinking part too.
I realized that the new functionality was already described in BRDs and FSDs. My existing RAG system already contained the baseline business processes.
One described what had changed. The other described how the application normally worked.
I just needed to combine them.
I had an email account created specifically for AI and automation work, so I used Power Automate and AI Builder to create three specialized roles:
- an Analysis Agent to interpret the request and identify the required operational knowledge;
- a Planning Agent to combine the new requirements with the retrieved business context;
- a Communication Agent to reply in the original email thread with the generated plan.
The product owner or business analyst could now send the testing request and supporting documents directly to the system’s dedicated email address. The system preserved the inputs, analyzed the change, retrieved the relevant workflows, generated an ordered plan, and replied with what it intended to test.
The Glue
The business-intelligence workflows ran in the Azure cloud (Power Automate).
The existing knowledge base and retrieval pipeline ran on an AWS EC2 server.
The GUI execution subsystem ran across Windows and WSL on my workstation.
All three environments now needed to cooperate.
I used SharePoint as a shared artifact workspace and Microsoft Graph API to access it from the server.
Each request received its own run folder containing:
00_Metadata
01_Input
02_Analysis
03_Planning
04_Execution
05_Evidence
06_Status
A server-side poller watched the run state and invoked the corresponding handler:
RECEIVED → Prepare the analysis input
ANALYZED → Retrieve the required operational knowledge
PLANNED → Retrieve and validate the execution plan
I initially tried using Graph change notifications. They refused to cooperate, even after I exposed an SSL-enabled endpoint. So I replaced the elegant solution with a polling loop and a processed-state tracker. It was less exciting. It also worked.
Did I Replace Myself?
Not completely.
The recorded execution completed approximately 15 consecutive GUI actions involving selections, text entry, navigation, and form progression.
The system captured screenshots before and after every action and compiled them into an execution replay.
But the complete 22-goal plan was not executed.
The server-to-workstation handoff remained manual.
The Atomic Execution Critic and Business Step Judge worked independently, but continuously invoking them made an already slow local execution loop even slower.
At the time, one multimodal action-selection request could take approximately one minute on my RTX 3060.
The final evidence package was also not automatically returned through the original email thread because the execution never reached the terminal state.
So I did not replace myself.
I built the eyes, brain, motor control, memory, communication channel, and workflow of my replacement.
Then I manually introduced some of them to each other.
When I demonstrated the prototype to my manager and CTO, they were genuinely impressed. Then they asked the inevitable question: How much would this cost?
My rough estimate for running the complete workflow with the GPT-5.4 computer-use model was higher than the cost of the existing six-person QA team (excluding me, of course). That ended the production discussion rather quickly.
Still, they understood that the system was far ahead of the organization’s immediate operational priorities. Like Howard Stark, I briefly felt limited by the technology of my time. Or, more accurately, by the technology budget of my company.
But the story still had a happy ending.
My manager never assigned me another QA task after that demonstration.
I was finally out of QA hell and free to spend my time building the systems I actually cared about.
Technically, the machine did not replace me. It convinced management to stop making me do the QA job. Close enough.
The Actual Lesson
The difficult part was never making a model click a button. The difficult part was everything around the click.
A useful system needed the request, domain knowledge, a plan, software access, current state, deterministic actuation, validation, evidence, and communication.
The VLM was valuable for semantic decisions.
Python geometry was better for coordinate resolution.
Durable artifacts and explicit states were better than hoping one massive prompt would remember the entire job.
I started this project because I hated repetitive GUI testing.
I ended up learning about multimodal inference, GUI grounding, retrieval systems, workflow automation, distributed coordination, state machines, and the boundaries between probabilistic reasoning and deterministic control.
That is a completely unreasonable amount of work to avoid a testing assignment.
I would probably do it again.
Technical Artifacts
The full technical report is available at: Reverse Engineering a Knowledge Worker
The architecture diagrams, sanitized prompts, workflow configurations, and representative implementation artifacts are available at: Project-REKNOW on GitHub.
