Martech Stack Evaluation
How to Evaluate Your Martech Stack
A martech stack evaluation reads your tools as a system — not a spreadsheet of line items. These are the checks the builder actually runs against your real stack, computed from a capability knowledge base and human-reviewed vendor judgments rather than filled in by hand.
No signup needed — tick your tools, get a scored evaluation in minutes.
Why most stacks are never properly evaluated
Stacks grow by accumulation. Without a structured evaluation, redundancy and wasted spend compound quietly.
<40%
of the platforms teams pay for are actually used — the rest is evaluation debt.
40-60%
is how much most teams underestimate real stack spend — licenses are only part of the cost.
~6%
of companies have AI fully embedded in their stack — most haven't evaluated readiness.
What the evaluation computes
Everyone else hands you a framework to fill in by hand. These 6 checks run against your real stack, and the engine — not the model — produces every number in them. The AI's job is to explain the findings, which is why it can cite them.
Capability overlap
Ranked overlap pairs, each with the shared capabilities and a reviewed judgment.Are two tools being paid for to do one job?
What it inspects:
- Every pair of tools in the stack, against a capability map of what each one does
- Human-reviewed vendor judgments — whether one genuinely substitutes for the other
- Which capabilities are core to a tool and which are supporting features
- Categories where running several tools is normal, so ad channels are not called redundant
How Stack Builder computes it: Every tool is mapped against a capability taxonomy of 128 capabilities, and overlapping pairs are ranked using human-reviewed vendor judgments — so "these two do the same thing" is a stored decision, not an opinion generated on the spot.
Redundancy and grade
A health score and an A–F grade, normalised by stack size.How much of this stack is duplicated, and how bad is that?
What it inspects:
- The severity-weighted share of tools that duplicate another tool
- Whether an overlap is a full substitution or a partial one
- Stack size, so a large stack is not penalised for having more possible pairs
How Stack Builder computes it: The grade comes from the severity-weighted share of tools that duplicate another, normalised by stack size. The AI never produces the number; it explains the one the engine computed.
Recoverable spend
Conservative, realistic and optimistic annual savings, plus the share of budget at risk.What is the duplication costing, and what could you get back?
What it inspects:
- Declared cost per tool where you have entered it
- Category spend benchmarks, scaled to company size, where you have not
- Which tool in an overlapping pair is the likelier one to drop
How Stack Builder computes it: Costs you have entered are used as-is; where you have not, category benchmarks scaled to company size fill the gap — and the report says which figures are declared and which are estimated.
Per-tool verdict
A keep / evaluate / remove call per tool, with the reason it was made.For each tool: keep it, look at it, or drop it?
What it inspects:
- Whether another tool in the stack covers this one’s core capabilities
- Lifecycle stage — a tool you have already marked sunsetting is read as leaving
- Whether the tool is connected to anything, and what would break without it
How Stack Builder computes it: Keep, evaluate or remove per tool, read against the rest of the stack — and lifecycle stages you have already set steer the call, so a tool you marked sunsetting is treated as leaving.
Topology
Hubs, orphans and single points of failure, named.Is the shape of the stack sound?
What it inspects:
- Tools everything depends on — a failure there takes the stack with it
- Tools connected to nothing, which are usually spend nobody is watching
- Whether data reaches the systems that need it
How Stack Builder computes it: The canvas already knows what connects to what, so hubs, orphans and single points of failure are read off the graph rather than inferred from a questionnaire.
Documentation maturity
A completeness score and the weakest dimension, so the verdict is caveated honestly.How much of this can the evaluation actually see?
What it inspects:
- How much of the stack carries ownership, cost, consent and integration metadata
- Which dimension is thinnest, because that is where the report is least confident
How Stack Builder computes it: Metadata coverage across ownership, cost, consent and integrations is scored, and the weakest dimension is named — the report tells you where it is least confident instead of pretending otherwise.
Stack DNA
The half of the evaluation that isn't a judgement call
The checks above are computed by the engine. Alongside them sits Stack DNA, which is computed — the same diagram in, the same numbers out, every time. It is what makes the overlap finding arguable in a budget meeting rather than dismissable as an opinion.
Connectivity, per tool
What fraction of the stack each tool actually touches, counted upstream and downstream. Hubs and orphans fall out of the arithmetic rather than the eye.
Redundancy groups
Tools with overlapping capabilities, grouped by category, with combined monthly cost and estimated saving. Computed from the capability taxonomy — not a guess about which logos look similar.
A 0-100 health score
One composite number for the stack, alongside a red/amber/green read on each tool with the reasons attached, so a score can always be traced back to what caused it.
The Solar System overlay
The same diagram rendered as a radial map, with the most-connected tools at the centre. A way to see your stack's gravity — which tools hold everything together, and which could quietly disappear.
Core (centre)
Identity, CDP, warehouse. Lose these and the stack stops.
Middle orbit
Activation channels, analytics, governance. Important, replaceable.
Outer orbit
Specialty tools, edge cases. Useful but optional.
Run the evaluation in four steps
Map your stack
Drag your tools from the 150+ vendor library onto the canvas and connect them to show how data flows.
Run the evaluation
The engine computes overlaps, the grade and the recoverable spend instantly; the AI narrative explains them, and the lenses (data flow, classification, consent) show the same findings on the canvas.
Read the scorecard
Get letter grades, redundancy and SPOF detection, recoverable spend, and specific recommendations per dimension.
Share and re-run
Export a presentation-ready report for the budget review, then re-evaluate quarterly off the same live diagram.
Which system am I looking at?
Four things on this site have names, and two of those names appear twice. Here is the whole map.
- Rings
- classify what a tool is. Five of them, from Scott Brinker’s model: System of Record, Engagement, Intelligence, Productivity, Integration.
- Perspectives
- choose what you are looking at. Four composite views — Architecture, Data & Privacy, Operations, Integration Health. Each bundles lenses and legends for one concern. They filter the canvas; they do not score anything.
- Dimensions
- score how well the stack performs. 6 engine checks that produce the A–F grade. This is the only one of the four that puts a number on your stack.
- Stack DNA
- is the evidence the dimensions are scored from. Computed connectivity metrics, redundancy groups and composite health values. You rarely read it directly — it is what the grade is built on.
Why “Architecture” and “Integration Health” appear twice. Each is both a Perspective and a scoring dimension. Same subject, different job: the Perspective filters the canvas so you can look at it, the dimension puts a grade on it. Seeing the name in both places means they are deliberately about the same concern — not that one is a copy of the other.
One more grade, unrelated to these. The free instant audit also returns an A–F letter, but it grades one thing only — how much of your stack is duplicated. It is not a short version of the 6-check evaluation, and the two letters will not match.
Looking for the views rather than the scores? See Perspectives and the 5 Rings.
Prefer it done for you?
Hand the evaluation to a martech architect
A two-week, fixed-price engagement: scored Stack DNA report, redundancy and SPOF analysis, and a 90-day action plan with owners.
See the done-for-you serviceThe method
Run the evaluation by hand
You do not need software to find what your stack is buying twice. You need an inventory, an honest list of what each tool actually does, and the discipline to judge overlapping pairs one at a time. Allow an afternoon for a 20-tool stack.
-
Inventory every tool, including the ones nobody owns
One row per tool. Columns: name, the job it does, who owns it, annual cost, and renewal date. Pull from the finance export rather than memory — the tools nobody remembers are exactly the ones the evaluation needs to catch.
-
Write down what each tool actually does
Against every row, list the capabilities you actually use it for — not the feature list on its website. Mark each one core (you bought the tool for this) or supporting (it happens to do it). The distinction decides every overlap below: two tools sharing a supporting capability is normal, two tools sharing a core one is a purchase you made twice.
-
Compare every pair and mark the real overlaps
Work through the pairs, not the categories — the expensive duplication is usually across category lines, where nobody is looking. For each pair sharing a capability, decide one thing: could one of these genuinely replace the other for that job? Write yes, no, or partly. "Partly" is the honest answer more often than either extreme, and it is the answer that survives a challenge.
-
Grade it on how much of the stack is duplicated
Count the tools that another tool could replace, weighting a full substitution heavier than a partial one, and divide by the number of tools you have. That share — not a rubric — is the grade. Divide by stack size or a big stack always looks worse than a small one, because pairs grow far faster than tools do.
-
Attach the money, then write the reason next to each pair
For every overlap, put the annual cost of the tool you would actually drop against it, and name the reason in one line. A number without a reason cannot be argued in a budget meeting or re-checked next quarter — and the reason is what tells you whether the cut is worth the migration.
Mid-market B2C retailer, 18 tools, one CDP, two overlapping email platforms.
| Tools with a full substitute in the stack | 3 of 18 |
|---|---|
| Tools with a partial substitute | 4 of 18 |
| Weighted redundancy | 28% of the stack |
| Grade | C |
| Annual cost on the droppable side | $96,000 |
| Largest single pair | Second email platform vs CDP messaging, $54,000 |
Reading it: The grade is driven by two email platforms and an analytics tool the CDP already covers — a consolidation problem, not a replatforming one. That is a very different budget request, and the pair-by-pair reasons are what make it survive the meeting.
The shortcut: Stack Builder does step 3 against a capability map and human-reviewed vendor judgments, so the pairs are computed rather than argued, and step 5 arrives with the costs already attached.
Frequently Asked Questions
What is a martech stack evaluation?
How is an evaluation different from a stack audit?
What framework should I use to evaluate my stack?
How often should I evaluate my martech stack?
Can I run an evaluation for free?
Can I evaluate a client's stack?
Evaluate your martech stack
Find duplicated capability, get an A–F grade and the spend attached to it, and walk into the budget review with evidence — not opinions. Free to start.
No credit card required. Free plan available.