William

William Β· Talent Sourcing Expert Β· October 4, 2026

How to Audit a Codebase Before You Acquire a Company, in Six Steps

Cover for a guide to auditing a codebase before an acquisition, beside a terminal showing a dependency installation log.

Key takeaways

  • β†’A code audit answers three questions: can we change this safely, can we run it, and what did the seller skip to hit a deadline.
  • β†’Ask for the written coding standard before you ask for the code. With no standard, every past review was an opinion, and that is finding number one.
  • β†’Tests, continuous integration and the dependency list show whether quality is enforced by machines or by two people's good habits.
  • β†’Security work deferred under launch pressure is invisible in a demo and becomes your liability the day you close.
  • β†’Every finding leaves the audit with a severity, an effort in engineering weeks and a price, or it changes nothing in the negotiation.

The product works in the demo. That is usually the only piece of a target company's engineering a buyer gets to see, and it is the piece that tells you the least. The codebase underneath decides what the next two years cost: how fast your team can ship on top of it, how many engineers you need just to keep it running, and whether a piece of security work the seller postponed becomes your incident to report. Most buyers of a small software company are not engineers, and the honest version of that problem is not that they cannot read code. It is that they do not know which questions separate an awkward codebase from an expensive one, so the judgment gets handed to whoever is available. What follows is a six-step audit a non-technical buyer can run with one independent senior engineer inside a normal diligence window, and finish with a list of findings that each carry a severity, an effort and a price.

1. Decide what the audit has to answer before you open the repo

An audit without a question list becomes a tour. Your reviewer reads code for a week, comes back with adjectives, and nothing in the deal changes. Write the questions first. Three of them carry almost all of the value:

  • Can we change this safely? How long does a small, real change take, and how confident is anyone that it broke nothing else.
  • Can we run this? What it takes to deploy, to recover from an outage, and to onboard an engineer who has never seen the system.
  • What was skipped to hit a deadline? Every shipped product has a list of things that were deferred. Nobody writes that list down, so you have to reconstruct it.

Then set the frame around those questions. Agree in writing which repositories, services and infrastructure accounts are in scope. Ask for read-only access rather than a guided walkthrough. Time-box the work, five to ten business days is realistic for a single-product startup, and name a deliverable: a findings list, not a report. Pick the reviewer yourself, and pick someone with no relationship to the seller. The common failure here is subtle: a buyer who cannot evaluate engineering also cannot evaluate the engineer evaluating the engineering, and the easiest person to borrow is the one the seller recommends.

Printed financial reports, a notepad and a calculator app on a phone laid out on a desk.
The code audit belongs in the same diligence file as the financials, with the same deliverable: numbers.

2. Read the code against a written standard, not against taste

"Clean code" is an opinion until there is something to compare against. So the first artifact to request is not the code, it is the standard: the style guide, the linter configuration, the architecture conventions the team says it follows. Large engineering organizations publish standards like these and keep revising them, because a standard that never changes is usually one nobody reads.

Both outcomes are useful. If a standard exists, the audit becomes a comparison: does the code match what the team says it does, and when did the document last change? If no standard exists, that is finding number one. It means every review in the company's history was a matter of individual preference, and it predicts most of what your reviewer is about to find.

Inside the code itself, a week of reading should cover:

  1. Structure and boundaries: can you tell what each module owns, or does changing a price break the email templates?
  2. Duplication and dead code: how many copies of the same logic exist, and how much of the repository is no longer reachable?
  3. The extremes: the largest files, the longest functions, the oldest untouched areas. Problems concentrate there.
  4. The README: can a new engineer run the project locally in a day, following only what is written?
  5. One real change: have your reviewer implement a small feature or fix in a branch and time it. Nothing reveals changeability like changing something.
A laptop screen showing a code editor with a file tree on the left and syntax-highlighted source on the right.
Structure and naming tell a reviewer more in an hour than a feature list does in a week.

3. Check what is automated: tests, pipeline and dependencies

A quality standard enforced by machines survives a bad quarter. One enforced by good intentions does not, and a bad quarter is exactly what an acquisition produces. So the automation layer is where you learn whether the quality you were shown is a property of the system or a property of the two people who built it.

On testing, the questions are blunt. Does an automated test suite exist? Does it pass on a clean checkout today? What does it actually cover, specifically on authentication, permissions, payments and anything that writes customer data? Automated tests are the cheapest quality control the seller could have bought, and when they are missing you are not buying a neutral absence, you are buying every bug that would have been caught in the years the suite did not exist.

On the pipeline: does continuous integration run on every change, does it block a merge when it fails, is deployment scripted and repeatable, how many manual steps stand between a merged change and production, and has anyone ever rolled back? Our walkthrough on how to set up a CI/CD pipeline covers what a working version of this looks like, which is the benchmark to hold the target against.

On dependencies: the install log is a map of risk you are inheriting. Count the third-party packages, check whether versions are pinned, look for packages that have had no release in years, run a vulnerability scan, and read the licenses against what you plan to do with the product after closing.

A terminal filling with dependency installation lines: cached and downloading wheel packages with sizes and transfer speeds.
Every third-party package in this log is risk you inherit on closing day, including the abandoned ones.

4. Find the security work the deadline pushed out

There is a pattern worth looking for by name. A team under launch pressure brings in a contractor to hit a date. The features ship on time, the security work slips, and it never returns to a roadmap because no ticket is ever written that says "we skipped this." The product demo looks identical either way. A startup that shipped fast and left user data exposed presents exactly like one that shipped fast and did not.

The checks that find it:

  • Secrets: API keys and passwords committed to the repository, including in the git history rather than only the current files.
  • Access: who holds production credentials today, including former contractors and employees who left.
  • Authorization: whether permission checks live in the API layer or only in the interface, which is the difference between a rule and a suggestion.
  • Customer data: what personal data is stored, where, encrypted or not, and who can query it directly.
  • Backups: not whether backups exist, but whether a restore has ever been performed and timed.
  • History: past incidents, what was done afterwards, and whether any regulatory regime applies to the data the company holds.

Price what you find as your liability, because after closing that is what it is. A breach involving customer data is one of the few findings in a code audit that can cost more than the company.

Lines of source code projected across an office wall, with a workstation cordoned off behind hazard tape.
The risks a product demo never shows: what the team deferred, and what nobody wrote down.

5. Judge the engineers you are buying, not only the repository

Code is a snapshot. The team is what keeps it alive after the wire transfer clears, so half of a useful audit is conversations rather than files.

Start with how a change reaches production. Who reviews it, is review mandatory, and how often does it happen? Peer review is how a standard actually spreads through a team, and the teams that treat it as a routine habit rather than a formality end up with a very different cost curve from the one where a single engineer merges to the main branch at midnight. Our guide on how to review a pull request describes the version worth asking about. Ask to see the review history itself, not a description of it.

Then ask what the company does to keep its engineers current: is there time budgeted for learning, does a junior engineer have a mentor, does anyone pair on hard changes? Training and mentorship are what pull deployment errors down and ship speed up over time, and they are also the first thing a seller quietly cuts in the year before a sale.

Finally, map the key-person risk. Ask every engineer the same question, which parts of the system only one person understands, and compare the answers. Then find out who intends to stay, what their compensation looks like against the current market, and what replacing them would cost you in time: our notes on how to onboard a remote developer are a reasonable proxy for how long a replacement takes to become productive.

A person at a desk watching two screens filled with charts and live data panels in a dark room.
Ask who runs this system the day after closing, and which parts only one person understands.

6. Turn every finding into a number before you negotiate

An audit that ends in adjectives changes nothing about a deal. An audit that ends in dollars changes the price. Every finding should leave the process with four fields: what it is, how severe it is, how much engineering time it takes to fix, and when it has to be done.

Sort severity into three buckets and resist the temptation to add a fourth. Blocking findings are the ones you will not close without resolving, typically exposed customer data or an unresolvable licensing problem. Ninety-day findings are the ones that make the first integration quarter miserable if left alone: no tests on the payment path, no reproducible deployment, secrets in the repository. Backlog findings are real but survivable, and they belong on the roadmap rather than in the negotiation.

Convert engineering weeks to money using your own loaded cost per engineer, not the seller's. Then decide how the total enters the transaction: a reduction in price, a holdback released when specific remediation is verified, or representations and warranties covering security and intellectual property. Any of the three is defensible. Noting the findings in an appendix and closing anyway is not.

One adjustment matters here. Technical debt compounds with scale. A system that is merely awkward with five engineers and a thousand users becomes a different class of problem at twenty engineers and fifty thousand, because the same missing test suite and the same tangled module now block more people more often. Price the trajectory you are buying, not only the state you inspected.

A printed cost sheet in close-up with dollar amounts on successive lines.
Each finding leaves the audit as a line item: severity, engineering weeks, and a number you can negotiate against.

Mistakes that turn a code audit into theater

  • Letting the seller choose the reviewer. Or borrowing one from a party that benefits from the deal closing. Independence is the whole value of the exercise.
  • Buying an opinion instead of a comparison. "The code is messy" is not actionable. "The code does not follow the standard the team published, in these four areas, costing roughly this much to correct" is.
  • Auditing the code and skipping the pipeline. Most of the recurring cost lives in the build, the deployment process and the dependency list, not in how the functions are named.
  • Treating the demo as evidence. A demo shows a prepared path on prepared data. It is the one part of the product guaranteed to work.
  • Reading "no bugs reported" as "no bugs". With no tests and no monitoring, the absence of reports is an absence of information.
  • Treating the audit as a one-off. Standards, reviews and testing are continuous practices. If the target has none of them, your first ninety days after closing have to install them, and that work belongs in the integration plan with a name attached.

Check what you would do in a diligence week

6 questions on the brief you just read. Pick one answer per question.

  1. 1. What is the first artifact to request before reading any code?

  2. 2. Your reviewer has five days. Which exercise tells you the most about how changeable the codebase is?

  3. 3. A target has no automated test suite. What does that mean for the deal?

  4. 4. Which security check is most often skipped by buyers?

  5. 5. The audit is finished. What makes it useful in the negotiation?

  6. 6. Why does technical debt get more expensive after the acquisition?

Score: 0 / 6

Run the six steps in order and the audit stops being a judgment call. You set the questions, you compare against a written standard, you check what the machines check, you reconstruct what the deadline pushed out, you talk to the people who will run it, and you leave with a number. If the deal is small, the whole thing fits in two weeks.

The one piece a buyer cannot improvise is the reviewer. If you do not have a senior engineer inside your own team who can do this independently, bring one in for the diligence window rather than borrowing one from the other side of the table. You can hire vetted software developers or a security engineer for exactly that kind of short, scoped engagement.

Frequently asked questions

How long should a code audit take before an acquisition?

Five to ten business days of one senior engineer's time is realistic for a single-product startup with one or two repositories. Multiple services, a mobile app and an infrastructure estate push it to three or four weeks. Scope it against what the deal needs rather than what is theoretically complete: the goal is a priced findings list, not an exhaustive review of every file.

Can I run a code audit if I am not technical myself?

You can run the process, but you cannot be the reviewer. Your part is setting the three questions, defining the scope and access, choosing a reviewer with no relationship to the seller, and translating findings into severity, effort and cost. The reviewer's part is the code. Keep those two jobs separate and the audit works even when the buyer has never written a line.

What if the target has no coding standard and no tests?

That is common in companies built fast, and it is not automatically a deal breaker. It is a cost: remediation work before you can ship confidently, and a slower first year while your team installs the practices that were never there. Estimate both in engineering weeks, convert to money at your own loaded cost, and bring that number into the price discussion rather than discovering it in month three.

What access do I need from the seller to do this properly?

Read-only access to the repositories including full git history, the continuous integration history, dependency manifests, an inventory of infrastructure and third-party services, the production access list, and any incident records. All of it under NDA, and it is normal to stage access so the most sensitive items arrive once the deal is reasonably advanced. A seller who refuses read-only repository access is giving you a finding of its own.

Ready to hire?

Vetted talent ready for US teams. No recruitment fees. Zero risk.

πŸ‡ΊπŸ‡Έ Trusted by companies across the United States