Transitioning Into AI Engineering: What to Learn and How to Actually Apply (Part 1 of 2)
In this lesson: Byron Mackay breaks down how to break into AI engineering right now: speed is no longer the signal and judgment is, planning and verification are now the core of the job, and agents, harnesses, and evals are the skills that show you can build with AI and get you hired.
Get reminders
One signup gets you join links, reminders, and the recording library. No ongoing commitment.
Top 3 takeaways
Speed is no longer the signal. Judgment is.
Agents can write code fast for anyone, so writing code fast doesn't set you apart anymore. Boilerplate, syntax memorization, line-by-line debugging, and hand-translating requirements into code are losing value. Reading code, system design, knowing your frameworks deeply, and understanding production code are gaining value. Because LLMs regress to the mean, the engineers who stand out are the ones who know what good looks like and steer the agent toward it.
Plan up front, verify after, and keep the work small
The new skills are precise specs, planning up front, verification and taste, and agent orchestration. The agent will build whatever you give it, so most of the effort moves to the plan before and the review after. Keep PRs small, skip review on simple low-risk changes, and scan for patterns instead of reading every line. Skip the review entirely and you build up verification debt. Once it ships, you own the code.
Agents, harnesses, and evals are what get you hired
Companies used to want engineers who knew about AI. Now they want engineers who are proficient with it. Byron points to three skills that show that: building and using agents, building harnesses around them, and evals. In his experience, evals alone can get you a job, because every team wants them and almost none have them.
Byron Mackay
Director of Learning, Gauntlet AI
Byron Mackay is a software engineer and the Director of Learning at Gauntlet AI. He describes his job as keeping up with AI and deciding what Gauntlet should teach. He spent years building mobile apps and learned Swift, Rails, and TypeScript, each because the next job needed it. This year alone he has seen more than 3,000 interviews through Gauntlet. This session is Part 1 of his two-part series on breaking into AI engineering.
Lesson notes
A written walkthrough of the lecture, covering how the engineering job market has changed, which skills are rising and falling, how to plan and review work done by coding agents, the skills that show AI proficiency, and how to stand out when you apply.
Getting your foot in the door
Byron opens with the numbers. The average job posting draws about 242 applications, and only around 3% of applicants get an interview. Cold outreach used to land a first interview roughly 15% of the time. Now it’s closer to 3%. Many resumes now read as AI-generated, so they all sound alike and recruiters have a harder time telling people apart. His advice is to get in through people. Every time he has had a referral, he has gotten an interview. He can’t point to a single cold application that clearly worked. When you do talk to a company, three things help most: a project you built specifically with that company in mind, a detail about the company you noticed through your own research, and plain language in your own voice.
What’s actually changing
The engineering role isn’t going away, Byron argues, but the work is shifting. At companies with deep AI adoption, simple junior roles are shrinking, while the pay premium for AI and architecture skills is growing. He has seen this in layoffs among friends: the people who lost their roles tended to be the ones without system design ability or the willingness to read code and direct the AI. He places this in a longer history. Compilers, structured programming, and object-oriented languages were each supposed to end engineering, and each one created more jobs. This shift is moving faster, which is why it feels unsteady. He uses Who Moved My Cheese to make the point: the people who keep looking for new skills are the ones who keep going. He also cites a talk by Werner Vogels, who argued that developers are entering a renaissance rather than becoming obsolete, as long as they evolve.
The skills map: declining, still essential, and new
Byron sorts engineering skills into three groups. Declining: writing boilerplate, memorizing syntax, line-by-line debugging, and hand-translating requirements into code. Still essential and rising in value: reading code, system design, knowing your frameworks, and understanding production code. System design gets special emphasis. In his experience, candidates who know it go much further in interviews, and Gauntlet built a Claude skill that teaches and quizzes system design fundamentals, which he plans to share with attendees. Knowing your frameworks matters more now because LLMs regress to the mean. His example is TypeScript’s discriminated unions: few developers use them, so an agent will never reach for one unless you ask for it explicitly. New skills to learn: precise specs and prompts, planning up front, verification and taste, and agent orchestration and context.
Planning is the core skill
For bigger work, Byron gets a few engineers in a room, turns on a voice recorder, and architects the solution out loud. The transcript becomes a PRD or write-up, the team reviews and edits it, then the agent turns it into tickets, which also get reviewed. For smaller but still significant features, he writes a spec with AI’s help, reviews it himself, and has another engineer review it. That sounds like a lot of work, and he admits it is. But the agent will write the code quickly either way. The planning is what makes sure it builds the right thing. He adds a caveat: plan in proportion to the task. Small fixes don’t need a big plan, new features need a short spec, and new systems get the full process.
Taste, verification, and owning the code
Taste means knowing what good looks like, and it comes from reading a lot of code, knowing your frameworks, and understanding your system. Byron points out that this was always what got engineers promoted. With smaller teams and faster agents, it is now becoming a requirement for everyone, and every engineer is starting to act like a CTO: strategizing, delegating, and verifying. Skipping verification creates verification debt. His example: a junior developer sent him a 300-line PR for what looked like a simple change. Byron had already pictured how he would solve it, couldn’t follow the PR, and told the developer it could probably be done in two lines. It could. Without that review, the codebase would have gained 300 lines of bloat. He echoes Werner Vogels’ point that code now shows up instantly, but understanding it still takes time.
How to review AI-generated code
Byron’s review habits: force the AI to work in small pieces, splitting a feature into separate database schema, service, API, and front-end changes instead of one huge PR. Keep PRs small enough to digest. He uses roughly 400 lines as a guide but says the number is arbitrary. Don’t review everything: if you trust the agent with a simple React hook or a basic CRUD service function, save your energy, because you are the bottleneck. And read for patterns. The best engineers he knows skim quickly and stop only when something looks out of place. Personal projects can move fast, but once code is in production at a real company, you own it and need to understand it.
Agents, harnesses, and evals
Four years ago, Byron got a job partly because he knew more about RAG than the next candidate. That edge is gone. Companies now want engineers who are proficient with AI, and three skills show it: agents, harnesses, and evals. A harness is the structure around an agent that gathers its context, brings a human into the loop, and keeps it from going off the rails. For agents, context is king, and orchestration is the key concept. Teams often build one agent after another until they need an orchestrator agent that hands out tasks, and he recommends adding an orchestrator only after you trust each agent to do its job alone. RAG isn’t dead, he says, but the word is overused. At heart it just means giving the agent sources of truth. On evals, Byron sees a leadership opportunity: product managers, not engineers, should write the evals, while the engineer sets up the eval framework and helps the team make it part of their workflow.
A side note on JEV
Byron also points attendees to JEV, from TypeSafe AI, and calls exploring it homework. It isn’t one of the three core skills, but he thinks it fills a long-missing piece in harness building. JEV is a general classifier. Instead of generating text, it returns each possible outcome with a confidence score, for example “yes, 60%; no, 40%.” That makes it much faster and cheaper than an LLM for classification tasks like sentiment analysis or choosing which tool to call. He describes it as a smart if statement: a way to classify, route, score, or branch where handwritten rules are too brittle.
Build for the company you want
Byron’s top tactic for landing a job is to build something for the company itself. Research the company and its product, find a gap, build a small version, and do it really well. Then bring it to an interview. It won’t work every time, but it shows you understand their product, their pain, and their customers, and he has seen it lead to strong offers. Contributing to a company’s open source project used to be another good route; he got to know Twitter engineers that way. That route is fading now that agents flood repositories with PRs.
Your assignment before Part 2
Pick a company you’d like to work for and spend no more than an hour researching its product. Then build something small using the habits from this session: plan it, write good specs, let the agent build, review the core functionality, make sure you understand the code, and generate a one-pager that maps the architecture. If you can, share it with someone, at lunch, on LinkedIn, or at a meetup. That part is optional. Part 2 covers the interview itself, drawing on what Byron has seen across thousands of interviews at Gauntlet.
Q&A highlights
Asked where to learn evals, Byron points to observability platforms, namely LangSmith, Langfuse, and his favorite, Braintrust, which built its platform around evals. He also suggests looking up error analysis, a method for figuring out what you should evaluate. Asked what to do if you’ve already shipped a production codebase you don’t fully understand, he says to start by having your coding agent build a one-pager with a graph of the high-level components and how data flows, then work down into specific areas one step at a time. He admits he’s in that spot himself on a project where he’s testing for product-market fit first. Asked about the best language to learn, his answer is whatever your next company uses. Asked about Kubernetes, he says it’s an important piece to understand. Everything is still relevant, and because LLMs regress to the mean, the extra knowledge is what keeps you above average.
FAQ
Is AI going to replace software engineers? +
Which engineering skills are losing value? +
Why does "LLMs regress to the mean" matter? +
How much planning should I do before letting an agent code? +
Do I need to review every line of AI-generated code? +
What skills show I'm proficient with AI? +
Where can I learn about evals? +
I already shipped code I don't fully understand. Where do I start? +
What programming language should I learn next? +
What's the best way to stand out when applying? +
What's next?
Keep building with the rest of Night School, or apply to Gauntlet — ten weeks of technical intensity with the best AI engineers we can find.