Most software selection processes measure the wrong things. Hourly rate, headcount, client logos, and the polish of a proposal are all easy to compare and none of them predict whether the software works. The things that do predict it are harder to see from outside: who actually does the work, how the firm behaves when an estimate turns out to be wrong, and whether anyone there is willing to tell you that the project you have described should not be built. This page covers the criteria worth weighting, the questions that expose them, and the situations where the right decision is to hire nobody.
Decide what kind of firm you need first
Three very different businesses present themselves as software partners, often using the same language. Choosing the wrong category is the most common and most expensive mistake in the process, because the mismatch does not show up until the engagement is already underway.
| Model | What you are buying | Works when | Fails when |
|---|---|---|---|
| Staffing arm | Individual engineers who work under your management and your process. | You have a technical leader, a defined backlog, and a genuine capacity gap. | Nobody internally owns the architecture, so five contractors produce five architectures. |
| Delivery partner | A team that owns a defined outcome end to end, including the decisions. | The problem crosses business and technical boundaries and needs someone accountable for the result. | You want to direct the work day to day but have not freed anyone up to do it. |
| Product team | Ongoing ownership of a product's direction, roadmap, and operation. | The software is the business, and it needs continuous investment rather than a project. | You needed one bounded system built and are now paying for a permanent function. |
Ask each firm directly which of these they are. A firm that claims to be all three at once is describing a sales strategy, not an operating model.
Find out who actually does the work
In a large consultancy, the people in the pitch are frequently a sales team and a senior partner who will bill a few hours a month. The engagement is then staffed with whoever is available, often the most junior people on the bench, because the economics of the model depend on it. This is not a scandal; it is how pyramid-shaped consulting is priced. It only becomes a problem when nobody tells you.
The fix is a specific question rather than a general one. Ask for the names of the people who will write the code and make architectural decisions, ask what else they are committed to during your engagement, and ask which of them is in the room right now. A firm that answers precisely is telling you the truth about its structure. A firm that answers with a capability statement is telling you something too.
- Who is the most senior person who will touch the code, and how many hours a week is that?
- What is the ratio of senior to junior engineers on this engagement, and who reviews whose work?
- If your best engineer leaves mid-engagement, what happens to our delivery date?
- Which parts of this will be subcontracted, and to whom?
- Who makes the call when a technical decision has a business tradeoff attached?
The questions that reveal outcomes
Evaluation questions are only useful if a weak firm cannot answer them well by accident. These are the ones where a rehearsed answer and an honest answer sound clearly different.
| Ask | A good answer sounds like | Warning sign |
|---|---|---|
| When did you last tell a client not to build something? | A specific example, including what they recommended instead and what it cost them in revenue. | "We always find a way to help." Nobody who bills for building will decline to build unless it is a habit. |
| How do you get from discovery to a number? | A small, fixed, defined first step, after which the next increment is priced against something real. | A confident fixed price before anyone has mapped the workflow, or an open-ended rate card with no first milestone. |
| What do we own at the end? | Code, documentation, infrastructure definitions, and accounts, with any reusable prior components named up front. | Ownership described as a license, or hosting and deployment that only they can access. |
| What is your handoff plan? | A described transfer: documentation, runbooks, and a period where your team operates the system with them watching. | Handoff treated as a phase to define later, or a support retainer offered as the answer. |
| How will we hear that something has gone wrong? | A named cadence, a named person, and an example of bad news delivered early on a previous engagement. | Status reporting that only travels through an account manager. |
| Which assumption, if wrong, breaks this estimate? | An immediate, specific answer: usually data quality, an integration they do not control, or a decision-maker's availability. | Reassurance that they have built this before, without naming a risk. |
| What would make you decline this engagement? | Real criteria: no internal decision-maker, a fixed date that cannot move, a problem that is organizational rather than technical. | Nothing. A firm with no disqualifying criteria has no filter, and you are inside it. |
| Who else should we be talking to? | Two or three named alternatives, including ones better suited to parts of the work. | A claim that nobody else does this. |
Ask about the exit at the start
Handoff terms are easy to agree before an engagement and nearly impossible to negotiate during one. By the time you want to leave, the leverage has moved. Raise it while you are still choosing, and treat the reaction as data. A firm that is comfortable with the question has designed for it, and a firm that finds it premature has designed for the opposite.
- Ownership of source code, and of anything built on top of pre-existing components
- Who holds the cloud accounts, domains, and production credentials during and after the work
- Whether infrastructure is defined in code you can read, or configured by hand in a console
- What documentation exists at the end, written for whom, and updated by whom
- Notice periods, and what a thirty-day wind-down actually produces
- Whether the system can be operated by a competent engineer who was not on the project
The practical test is simple: if this firm disappeared tomorrow, could a new team pick the system up? If the honest answer is no, you have not bought software. You have bought a dependency.
Check how they surface bad news
Every non-trivial engagement produces bad news. An integration is undocumented, the data is worse than anyone said, a requirement that seemed peripheral turns out to reach into everything. The difference between a recoverable project and a failed one is almost never whether bad news occurred. It is how many weeks passed before anyone said it out loud.
Firms that report through an account manager tend to smooth things, because their job is the relationship. Ask for direct contact between the people doing the work and the people who own the outcome, and ask how scope changes get raised. Incremental delivery matters here for a structural reason rather than a methodological one: when the next increment is agreed rather than assumed, a mismatch surfaces as a decision you make in month two instead of a surprise you absorb in month fourteen.
References you choose, not references they choose
A curated reference tells you the firm has at least one happy client. That is a low bar. Ask instead for a list of every client from the past two years and pick from it yourself, including any engagement that ended early. Reluctance at this step is itself informative.
- 1
Ask for a full client list rather than three names, and choose the calls yourself.
- 2
Ask each reference what went wrong, and how you found out about it.
- 3
Ask who was actually on the team, and whether those people matched the pitch.
- 4
Ask what happened after launch: who operates the system now, and how hard was that transition.
- 5
Ask whether they would hire the firm again for the same work, and for different work.
When not to hire an outside partner at all
The most useful outcome of an evaluation is sometimes the decision not to run it. Four situations come up repeatedly, and in all four an outside partner makes things worse rather than better.
- The problem is organizational, not technical. Two departments disagree about who owns a process, or an incentive structure rewards the behaviour you are trying to eliminate. Software built on top of that disagreement encodes it and makes it harder to change later.
- You have a capable internal team and only need capacity. Bring in contractors under your own technical leadership. Hiring a delivery partner to work around your own engineers creates two architectures and a political problem.
- The requirement is genuinely commodity. Payroll, general ledger, email, help desk, and standard e-commerce are solved. If your process is unusual only because of an accident of history, change the process and buy the product.
- Nobody internally can be freed up to make decisions. Discovery-heavy work consumes a real decision-maker's attention for a real number of hours each week. Without that, the partner will guess, and you will pay to rebuild the guesses.
None of this argues against outside help. It argues that the value of a partner comes from ownership and judgement, and both require a problem that is actually technical and a client who is actually available. When those two conditions hold, the selection criteria above matter more than price. When they do not, no partner selection process will save the project, and the cheapest thing you can do is find that out before signing.
Common questions
- How many firms should we shortlist?
- Three is usually enough, and more than four tends to degrade the process rather than improve it. Depth of evaluation predicts outcomes better than breadth, and a long shortlist pushes you toward comparing proposals on the shallow attributes that are easy to line up in a spreadsheet. Spend the saved time on reference calls you chose yourself.
- Should we run a formal RFP?
- Only if procurement requires it. An RFP forces every firm to price the same written scope, which rewards whoever is most willing to assume that the scope is correct. For discovery-heavy operational software the scope is usually wrong in ways nobody has found yet, so you end up selecting for optimism. A short paid assessment from two firms tells you more than ten RFP responses.
- Is a bigger firm a safer choice?
- It is safer politically and not necessarily safer operationally. Large firms bring process, continuity, and the ability to absorb a departure, but they also staff in a pyramid, which means the senior people you met are unlikely to do the work. Small firms carry key-person risk that you should ask about directly. Both risks are manageable once they are named.
- How much should the first engagement cost?
- Small enough that you can walk away from it without regret. A focused assessment or a clearly bounded first release is the point: it produces something usable, and it tells both sides whether the working relationship is any good before either has committed to a large number. Treat a partner's willingness to start small as a selection criterion in its own right.
- What if we cannot tell the difference between two good firms?
- Weight the people who will actually be on the engagement over the firm's track record, since delivery is done by individuals rather than by brands. If that is still a tie, choose based on how each firm handled the hardest question you asked and how quickly they named a risk without being pressed.
- How do we know an engagement is going badly early enough to act?
- Working software in front of real users is the only reliable signal; documents, status decks, and demos on synthetic data are not. If the first sixty days do not put something into production or produce a decision you can act on, that is the moment to raise it rather than wait for the next milestone. Say it directly. Most problems at that stage are scope, sequencing, or communication issues that can be corrected in the next increment.