Skip to content
NEW: Contract & SLA Management is now in open beta. Learn more →

Martech Stack Evaluation

How to Evaluate Your Martech Stack

A martech stack evaluation reads your tools as a system — not a spreadsheet of line items. These are the checks the builder actually runs against your real stack, computed from a capability knowledge base and human-reviewed vendor judgments rather than filled in by hand.

No signup needed — tick your tools, get a scored evaluation in minutes.

Why most stacks are never properly evaluated

Stacks grow by accumulation. Without a structured evaluation, redundancy and wasted spend compound quietly.

<40%

of the platforms teams pay for are actually used — the rest is evaluation debt.

40-60%

is how much most teams underestimate real stack spend — licenses are only part of the cost.

~6%

of companies have AI fully embedded in their stack — most haven't evaluated readiness.

What the evaluation computes

Everyone else hands you a framework to fill in by hand. These 6 checks run against your real stack, and the engine — not the model — produces every number in them. The AI's job is to explain the findings, which is why it can cite them.

1

Capability overlap

Ranked overlap pairs, each with the shared capabilities and a reviewed judgment.

Are two tools being paid for to do one job?

What it inspects:

  • Every pair of tools in the stack, against a capability map of what each one does
  • Human-reviewed vendor judgments — whether one genuinely substitutes for the other
  • Which capabilities are core to a tool and which are supporting features
  • Categories where running several tools is normal, so ad channels are not called redundant

How Stack Builder computes it: Every tool is mapped against a capability taxonomy of 128 capabilities, and overlapping pairs are ranked using human-reviewed vendor judgments — so "these two do the same thing" is a stored decision, not an opinion generated on the spot.

2

Redundancy and grade

A health score and an A–F grade, normalised by stack size.

How much of this stack is duplicated, and how bad is that?

What it inspects:

  • The severity-weighted share of tools that duplicate another tool
  • Whether an overlap is a full substitution or a partial one
  • Stack size, so a large stack is not penalised for having more possible pairs

How Stack Builder computes it: The grade comes from the severity-weighted share of tools that duplicate another, normalised by stack size. The AI never produces the number; it explains the one the engine computed.

3

Recoverable spend

Conservative, realistic and optimistic annual savings, plus the share of budget at risk.

What is the duplication costing, and what could you get back?

What it inspects:

  • Declared cost per tool where you have entered it
  • Category spend benchmarks, scaled to company size, where you have not
  • Which tool in an overlapping pair is the likelier one to drop

How Stack Builder computes it: Costs you have entered are used as-is; where you have not, category benchmarks scaled to company size fill the gap — and the report says which figures are declared and which are estimated.

4

Per-tool verdict

A keep / evaluate / remove call per tool, with the reason it was made.

For each tool: keep it, look at it, or drop it?

What it inspects:

  • Whether another tool in the stack covers this one’s core capabilities
  • Lifecycle stage — a tool you have already marked sunsetting is read as leaving
  • Whether the tool is connected to anything, and what would break without it

How Stack Builder computes it: Keep, evaluate or remove per tool, read against the rest of the stack — and lifecycle stages you have already set steer the call, so a tool you marked sunsetting is treated as leaving.

5

Topology

Hubs, orphans and single points of failure, named.

Is the shape of the stack sound?

What it inspects:

  • Tools everything depends on — a failure there takes the stack with it
  • Tools connected to nothing, which are usually spend nobody is watching
  • Whether data reaches the systems that need it

How Stack Builder computes it: The canvas already knows what connects to what, so hubs, orphans and single points of failure are read off the graph rather than inferred from a questionnaire.

6

Documentation maturity

A completeness score and the weakest dimension, so the verdict is caveated honestly.

How much of this can the evaluation actually see?

What it inspects:

  • How much of the stack carries ownership, cost, consent and integration metadata
  • Which dimension is thinnest, because that is where the report is least confident

How Stack Builder computes it: Metadata coverage across ownership, cost, consent and integrations is scored, and the weakest dimension is named — the report tells you where it is least confident instead of pretending otherwise.

Stack DNA

The half of the evaluation that isn't a judgement call

The checks above are computed by the engine. Alongside them sits Stack DNA, which is computed — the same diagram in, the same numbers out, every time. It is what makes the overlap finding arguable in a budget meeting rather than dismissable as an opinion.

Connectivity, per tool

What fraction of the stack each tool actually touches, counted upstream and downstream. Hubs and orphans fall out of the arithmetic rather than the eye.

Redundancy groups

Tools with overlapping capabilities, grouped by category, with combined monthly cost and estimated saving. Computed from the capability taxonomy — not a guess about which logos look similar.

A 0-100 health score

One composite number for the stack, alongside a red/amber/green read on each tool with the reasons attached, so a score can always be traced back to what caused it.

The Solar System overlay

The same diagram rendered as a radial map, with the most-connected tools at the centre. A way to see your stack's gravity — which tools hold everything together, and which could quietly disappear.

☀️

Core (centre)

Identity, CDP, warehouse. Lose these and the stack stops.

🪐

Middle orbit

Activation channels, analytics, governance. Important, replaceable.

🌑

Outer orbit

Specialty tools, edge cases. Useful but optional.

Run the evaluation in four steps

1

Map your stack

Drag your tools from the 150+ vendor library onto the canvas and connect them to show how data flows.

2

Run the evaluation

The engine computes overlaps, the grade and the recoverable spend instantly; the AI narrative explains them, and the lenses (data flow, classification, consent) show the same findings on the canvas.

3

Read the scorecard

Get letter grades, redundancy and SPOF detection, recoverable spend, and specific recommendations per dimension.

4

Share and re-run

Export a presentation-ready report for the budget review, then re-evaluate quarterly off the same live diagram.

Which system am I looking at?

Four things on this site have names, and two of those names appear twice. Here is the whole map.

Rings
classify what a tool is. Five of them, from Scott Brinker’s model: System of Record, Engagement, Intelligence, Productivity, Integration.
Perspectives
choose what you are looking at. Four composite views — Architecture, Data & Privacy, Operations, Integration Health. Each bundles lenses and legends for one concern. They filter the canvas; they do not score anything.
Dimensions
score how well the stack performs. 6 engine checks that produce the A–F grade. This is the only one of the four that puts a number on your stack.
Stack DNA
is the evidence the dimensions are scored from. Computed connectivity metrics, redundancy groups and composite health values. You rarely read it directly — it is what the grade is built on.

Why “Architecture” and “Integration Health” appear twice. Each is both a Perspective and a scoring dimension. Same subject, different job: the Perspective filters the canvas so you can look at it, the dimension puts a grade on it. Seeing the name in both places means they are deliberately about the same concern — not that one is a copy of the other.

One more grade, unrelated to these. The free instant audit also returns an A–F letter, but it grades one thing only — how much of your stack is duplicated. It is not a short version of the 6-check evaluation, and the two letters will not match.

Looking for the views rather than the scores? See Perspectives and the 5 Rings.

Prefer it done for you?

Hand the evaluation to a martech architect

A two-week, fixed-price engagement: scored Stack DNA report, redundancy and SPOF analysis, and a 90-day action plan with owners.

See the done-for-you service

The method

Run the evaluation by hand

You do not need software to find what your stack is buying twice. You need an inventory, an honest list of what each tool actually does, and the discipline to judge overlapping pairs one at a time. Allow an afternoon for a 20-tool stack.

  1. Inventory every tool, including the ones nobody owns

    One row per tool. Columns: name, the job it does, who owns it, annual cost, and renewal date. Pull from the finance export rather than memory — the tools nobody remembers are exactly the ones the evaluation needs to catch.

  2. Write down what each tool actually does

    Against every row, list the capabilities you actually use it for — not the feature list on its website. Mark each one core (you bought the tool for this) or supporting (it happens to do it). The distinction decides every overlap below: two tools sharing a supporting capability is normal, two tools sharing a core one is a purchase you made twice.

  3. Compare every pair and mark the real overlaps

    Work through the pairs, not the categories — the expensive duplication is usually across category lines, where nobody is looking. For each pair sharing a capability, decide one thing: could one of these genuinely replace the other for that job? Write yes, no, or partly. "Partly" is the honest answer more often than either extreme, and it is the answer that survives a challenge.

  4. Grade it on how much of the stack is duplicated

    Count the tools that another tool could replace, weighting a full substitution heavier than a partial one, and divide by the number of tools you have. That share — not a rubric — is the grade. Divide by stack size or a big stack always looks worse than a small one, because pairs grow far faster than tools do.

  5. Attach the money, then write the reason next to each pair

    For every overlap, put the annual cost of the tool you would actually drop against it, and name the reason in one line. A number without a reason cannot be argued in a budget meeting or re-checked next quarter — and the reason is what tells you whether the cut is worth the migration.

Worked example Illustrative

Mid-market B2C retailer, 18 tools, one CDP, two overlapping email platforms.

Tools with a full substitute in the stack 3 of 18
Tools with a partial substitute 4 of 18
Weighted redundancy 28% of the stack
Grade C
Annual cost on the droppable side $96,000
Largest single pair Second email platform vs CDP messaging, $54,000

Reading it: The grade is driven by two email platforms and an analytics tool the CDP already covers — a consolidation problem, not a replatforming one. That is a very different budget request, and the pair-by-pair reasons are what make it survive the meeting.

The shortcut: Stack Builder does step 3 against a capability map and human-reviewed vendor judgments, so the pairs are computed rather than argued, and step 5 arrives with the costs already attached.

Frequently Asked Questions

What is a martech stack evaluation?
A martech stack evaluation is a structured review of your marketing technology as a system rather than tool by tool. Stack Builder computes it: every pair of tools is compared against a capability map to find duplicated capability, the duplication is scored into an A–F grade, and the money attached to it is estimated from your costs or category benchmarks. You get overlaps, a grade, recoverable spend, a keep/evaluate/remove call per tool, and the topology risks — evidence you can defend in a budget review.
How is an evaluation different from a stack audit?
They overlap. An audit tends to focus on inventory and findings ("what do we have, what is wrong"). An evaluation goes on to a decision ("how good is this, and what should change"). Stack Builder does both from the same diagram and the same engine — see the Martech Stack Audit page for the audit-first workflow.
What framework should I use to evaluate my stack?
Draw the stack, then let the engine check it: capability overlap between every pair of tools, a redundancy grade, the spend attached to the duplication, a verdict per tool, the topology risks, and how much of the stack is documented well enough to judge. The point is that the findings are computed from a capability knowledge base and reviewed vendor judgments — the AI explains them, it does not invent them.
How often should I evaluate my martech stack?
Quarterly is a healthy cadence — frequent enough to catch stack sprawl before it compounds, infrequent enough to act between reviews. Because the evaluation runs off a live diagram, a re-evaluation is minutes, not weeks.
Can I run an evaluation for free?
Yes. The free Instant Stack Audit lets you tick your tools and see overlaps, recoverable spend, and a grade with no signup. The full builder (free plan: 3 stacks) adds the lenses and the AI evaluation behind the full report.
Can I evaluate a client's stack?
Yes. Consultants use Stack Builder to evaluate client stacks and hand over scored, presentation-ready reports. The Consultant plan adds unlimited saves and white-labeled exports.

Evaluate your martech stack

Find duplicated capability, get an A–F grade and the spend attached to it, and walk into the budget review with evidence — not opinions. Free to start.

No credit card required. Free plan available.