
Software Engineer, Verification Fleet
Product.ai
Job description
Build the checkout robots that prove a code works. Real browsers driving real carts across hundreds of thousands of stores.
Product.ai is the verified truth layer for shopping. When a person or an AI agent needs to know what is actually true about a purchase, we answer with proof. SimplyCodes is our first proof at scale, the code verification service that shows shoppers the codes that actually work instead of a wall of dead ones. It earns about $22 million a year at roughly 60% margins. We are 100% founder-owned, profitable, and bootstrapped since 2009. No outside investors, no board. Fewer than twenty operators, outbuilding companies 10x our size.
Why This Role Exists
When SimplyCodes tells a shopper a code works, a machine should have proved it. A checkout robot went to the store, added an item, applied the code, and watched what happened at the cart. You own that fleet.
Human spot-checking cannot reach the long tail, so the fleet is now the primary producer of the verified checkout evidence this company sells. Today the robots reach only a fraction of the stores beyond the big standardized platforms. Your job is to multiply that reach across hundreds of thousands of stores. Every store you add turns a guess into a verified claim. That $22 million is the engine that funds everything here, and those claims keep the merchant pages behind it worth ranking and trusting. The same proof is now a paid product for machines. Anonymous API access closed August 15; legacy free codes retire September 30.
One boundary. This seat proves the code works, with robots at real carts. A sibling posting owns the commerce-data supply chain, meaning what enters the system and how fresh it stays. If deciding what we know pulls you harder than proving it true, apply there.
Agents write most of the code here, so the scarce thing is judgment, the design taste that keeps a fleet correct and cheap while it multiplies. You decide what to build and how you will prove it holds, with a high technical bar underneath.
The System You'll Need to Model
-
Browser automation against a thousand different carts. Every e-commerce platform breaks differently. The code field hides behind a different click on Shopify, Magento, WooCommerce, BigCommerce, and a long tail of custom storefronts. The leverage craft is classification. Collapse hundreds of thousands of merchants into a small set of platform families, and one automation recipe covers thousands of stores instead of one.
-
Fleet economics. Every large crawler answers to the same unit-economics law. A check that costs more to run than the commission it protects is a loss. You are always trading how often you re-test against what the test is worth. The fleet has to earn its own keep, store by store.
-
The coverage-accuracy frontier. Pushing coverage outward means automating messier, stranger stores, where a robot is likelier to misread the cart. Reach and certainty pull against each other, the same precision-recall tension every verification system faces. Keeping both as you scale is the job.
-
Evasion versus detection. Stores and their anti-bot vendors do not want to be automated. Fingerprint surface, rate limits, and challenge walls move constantly, and headless detection improves every quarter. You navigate that contest at scale without breaking the store or the law.
-
Verdicts need tests; tests need verdicts. You build beside the verification-science seat we are hiring now, which owns the machine verdicts that decide what we claim is true. Those verdicts are only as good as the checkout evidence your fleet produces, and your fleet only knows where to test because that scoring shows where the truth is thin. Two seats, one loop.
-
Cortex, the brain you build inside. You work inside Cortex, the shared AI brain that runs the company and the product family we sell; it answers its own questions from more than 8,600 documents. The company moves at that brain's speed; no spec stays current for a quarter. You build inside the thing we sell rather than using AI on the side.
If reading that energizes you, keep going. If it feels overwhelming or underspecified, this isn't the right fit.
What You Will Own
-
The fleet platform, end to end
. Browser automation at fleet scale is the system class behind every serious crawler, price-intelligence engine, and synthetic-monitoring product; here it points at real checkouts. Yours spans the automation layer and the platform-family classifier, plus the scheduler that decides which stores to test and when. Agents write much of the code; you own the design, the failure modes, and the verdict on what ships.
-
Fleet economics
. Cost per verified checkout, held below the commission each check protects. You run the fleet's spend the way a strong platform team runs infrastructure, as a business number you can defend rather than a bill you discover.
-
The instrumentation that proves it
. Fleet observability and evidence capture. Dashboards and ledgers show, for any claim, when it was last tested, whether the robot made it to the cart, and what the check cost. Correctness you can watch instead of assert.
-
The number, co-signed
. Within your first quarter you co-sign a seat charter. It names one machine-checkable number that proves the seat works and writes down what you decide freely versus what you propose for the founder to sign. You own a number, not a backlog.
Who You Are
You reason in invariants, failure modes, and tradeoffs. Handed a checkout flow you have never seen, you can sketch the three ways it will break before you write a line. You see the platform family behind a one-off store, and the shared recipe behind a hundred one-off stores. When a robot fails at 2 a.m., your first question is structural: what class of store did we just discover?
You move fluidly between architecture and shipped code; a classifier design in the morning can be a deployed test by night. You treat agents as leverage you verify, not autocomplete you trust, which means you can point at a system you shipped, name the hardest failure you personally diagnosed in it, and say what you changed. You can do this job by hand and prove it, and that mastery is what lets you trust, or reject, what an agent hands back. A redo cycle costs far more here than compute does.
You have built browser automation, web scraping, or large-scale crawling systems and operated them in production. You know what a fleet of headless browsers does to your infrastructure bill and your on-call sleep. You have reverse-engineered a site that did not want to be automated, and won. Production scars in that craft are the strongest accelerant here; what we test first is whether you can model a system you have never seen and prove your fix holds. The scars can come from price intelligence, ad verification, search crawling, synthetic monitoring, or test automation at real scale; the mechanics are the same. Playwright, Puppeteer, headless Chrome, proxy rotation, and queue-backed job systems are familiar ground; Node.js and Python are daily tools. Here you layer on the next decade of the craft. You direct coding agents and verify what they return. You build evidence systems beside an LLM-evaluation seat and run a fleet's economics like a P&L. We care about the artifact and the reasoning far more than where you did it. No degree to check.
Who this isn't for.
This is wrong if you guard a single lane and call the rest someone else's department. You own the fleet across automation, classification, infrastructure, and cost, and "that's not my job" ends the conversation. It's wrong if you pick technologies for how they'll look on your next resume rather than for what the fleet needs tonight. Wrong if you optimize for influence over output and dress hard news in stakeholder euphemism. Wrong if you wait to be told what to test instead of reading the system and deciding. And wrong if your code is whatever the model handed you and you couldn't say why it's right, or if you're comfortable letting an agent grade its own work. You'll be happiest here if your idea of craft is a fleet of robots that proves, store after store, that a code is real.
How We Evaluate
We don't run traditional engineering interviews.
-
Async video screen. Brief and on your own time — about fifteen minutes. We want to see how you think, not how you present.
-
Calls with company stakeholders. Short conversations with the people you'd build beside.
-
Conversation with the founder. How you reason about coverage, cost, and truth at fleet scale, and where you push back.
-
Paid work trial. A paid four-day engineering trial — real work, in our real environment, shipping to our real platform. We watch how you get grounded in the system, whether you write the spec before the build, how you verify what your agents produce, and whether your self-assessment is honest. We both learn more in four days than in forty hours of interviews.
If the work above reads like yours but your resume is unconventional, apply anyway. We hire on the work and the reasoning, not the pedigree.
Compensation & Ownership
Total first-year comp: $380,000 – $475,000
— base, plus performance-based ownership and profit-share programs.
Base: $250,000 – $310,000
, top of market for senior engineering.
Eligibility for the company's ownership and profit-share programs — grants are performance-based, with terms discussed at the offer stage. We cover 100% of family insurance premiums. Your token budget is effectively unlimited, steered by return, never capped. The model is built to mint partners.
Based in Santa Monica, Los Angeles — in person, five days a week. The rooms are real rooms. Relocation support available for the right builder.