How we score tools and skills on EcomAgentTools — and why most review sites won't show you this

Maxwell Trent · Product Researcher · EcomAgentTools

A rubric card showing weighted scoring axes with source labels

See exactly how EcomAgentTools scores tools and skills: our weighted rubric, community rating adaptive weights, source transparency, and the one rule that prevents editorial bias.

How to use this report

Start with the evidence scope and test conditions, then compare the candidates against your own platform, workload, and approval rules. Prices, features, and platform support reflect the review date in the article; check the official page again before buying or connecting store data.

I spent most of a week reading the methodology pages behind review-site scores: G2, Capterra, Enjyn, and a dozen others.

The problem appeared quickly. Almost nobody shows the complete calculation. G2 comes closest: it explains that "Satisfaction" and "Market Presence" both matter and that older reviews lose weight through exponential decay. The exact weights and formula remain unpublished.

Without the formula, an operator can't tell whether a high score means users like the product or the vendor simply has a large team and marketing presence. It also cannot tell you how hard the product will be to operate; that is why we publish a separate difficulty-level system.

So this article shows the EcomAgentTools calculation in full: weights, sources, missing-data treatment, and the rule that stops us making exceptions for products we happen to like.

The short version

Every tool and skill receives a score out of 100 from a fixed weighted rubric. Each axis names its source, so readers can see whether the number comes from community ratings, editorial assessment, or derived fields such as setup time and platform coverage.

Prompts don't receive scores. A few lines of text don't become more useful because we print 82 beside them.

Where the idea came from

Enjyn's methodology provided a useful starting point. It normalizes public G2 and Capterra data to a 10-point scale, then applies the same five weighted axes across its catalogue. When the team has used a product in production, first-hand data can replace an individual axis and is labelled "operator"; otherwise the public review data remains.

The idea worth borrowing was fixed weights, visible sources, and clearly labelled operator evidence.

Our catalogue needed different missing-data handling. It contains many startup and scale-up ecommerce tools, and 68% had no G2 reviews at all. A G2-dependent model would effectively abandon two-thirds of the catalogue.

Community data can still become the strongest signal, but only after it exists. Until then, the remaining evidence carries the weight instead of treating missing ratings as a zero.

The Tool rubric

Five axes. Fixed across all 83 tools.

| Axis | Weight | What it measures | Where the data comes from | |------|--------|-----------------|--------------------------| | User ratings | 30% | EcomAgentTools community reviews | Community (your actual ratings) | | Capability | 25% | How good the tool is at its core job | Editorial assessment | | Time to value | 15% | How fast you get a useful result | Derived from our guidance data | | Price to value | 15% | What you get for what you pay | Editorial assessment | | Market presence | 15% | How active the tool's ecosystem is | Derived from G2 reviews, PH upvotes, GitHub stars |

Three details matter when reading this table.

First: user ratings carry the most weight, but they start at zero. If nobody has rated a tool yet — and right now, that's most of them — the community rating axis simply gets excluded and the remaining four axes redistribute the weight proportionally. The moment someone leaves a rating, the community axis wakes up. But it doesn't jump straight to 30%. It ramps up gradually: 1-2 ratings gives it about 6%, 3-9 gives it around 18%, and at 10+ ratings it reaches the full 30%.

This prevents one angry reviewer from controlling a score on a small site. Early results are therefore mostly editorial; community evidence takes over gradually as the sample grows. That is deliberate, and the source labels should make it visible.

Second: market presence is a proxy, not the result we really want. The current inputs are G2 review counts, Product Hunt upvotes, and GitHub stars because traffic data such as Ahrefs is not yet part of the automated pipeline. A product with 1,000+ G2 reviews and 500+ PH upvotes has more visible market evidence than one with neither, but the measure is still indirect. If Ahrefs is added, it will replace this axis rather than being stacked on top.

Third: price-to-value is editorial, but it follows a written rule. Free tools receive 9, products with 4+ pricing tiers receive 7.5, those with 2-3 tiers receive 6, and a single opaque "contact sales" option receives 5. The rule is rough. At least it is applied consistently and can be challenged openly.

The Skill rubric

Also five axes. Community weight is even higher here because skills live and die by whether people can actually use them.

| Axis | Weight | What it measures | |------|--------|-----------------| | User ratings | 35% | EcomAgentTools community reviews | | Capability | 25% | Effectiveness + ease of use + time to value | | Documentation | 15% | How well the skill is documented | | Reproducibility | 15% | Can you follow the docs and get it working? | | Ecommerce fit | 10% | How specifically this solves ecommerce problems |

Same adaptive logic applies: community weight scales from 0% to 35% as ratings accumulate.

Reproducibility matters because a polished README does not prove the installation works. For now, this axis still uses time-to-value as an editorial proxy. The better evidence will be real reproduction records: "I followed the guide and it worked" versus "step 3 returned a 500 error."

Why no scores for Prompts?

Prompts are usually only a few sentences. Giving one 78 and another 82 would suggest a precision the evidence cannot support. We now order prompts through editorial curation without displaying a score. The old descending numbers—94, 92, 90—were decorative, so we removed them. A ranking should look like a ranking.

The one rule that matters

Before I built any of this, I wrote down one constraint and taped it to my monitor: same rubric for every item, no exceptions.

The rule matters because scoring 83 tools by hand creates plenty of opportunities to bend the model. One company has a good demo, another has a founder you like, and a third simply feels more established. If the rubric allows convenient exceptions, people will use them. No corruption required—ordinary human preference is enough.

A fixed rubric makes Shopify Magic and an inventory planner with 12 users pass through the same calculation. If a result looks wrong, we review the weights or the evidence. We don't quietly override one product.

This is also why we publish the weights. If you disagree with them — if you think market presence should be higher, or capability should be lower — you can recalculate any tool's score yourself. The data is all there.

What's still broken

The model still has real gaps. They need to stay visible.

Community ratings are currently empty. The adaptive weighting system is built and deployed, but it's waiting for actual ratings. Until then, all scores are editorial. If you've used any of the tools or skills on the site, leaving a rating is genuinely the most impactful thing you can do.

Market presence is the wrong proxy. G2 review counts and PH upvotes correlate with market activity, but they're lagging indicators. A tool could be growing fast with zero G2 presence because its user base doesn't overlap with G2's demographic. The right metric is website traffic, and we need to integrate that.

Skills need real reproducibility data. Right now the reproducibility score is a proxy derived from other ratings. The only way to fix this is to actually attempt to reproduce skills and document the results. That's labor-intensive and we haven't built the pipeline for it yet.

The 100-point scale is compressed. Most tools cluster between 78 and 89. That's because several axes (time-to-value, market presence) produce similar scores across many tools. A wider distribution would be more useful for differentiation, but widening it artificially would mean inventing differences that don't exist. I'd rather have honest clustering than fake precision.

What we'll fix next

The roadmap, in order:

  1. Get real ratings in the system. Everything else is theoretical until people actually rate things.
  2. Integrate Ahrefs data for market presence. Replace the current proxy with actual website traffic metrics.
  3. Build a skill reproduction pipeline. Track which skills have been verified as reproducible and by whom.
  4. Publish a public methodology page. This blog post is the first draft of that page. Once we've stabilized the system, the weights and logic will live at /about/scoring with per-axis documentation.

The meta-point

I didn't build this because I think numbers solve everything. I built it because most review sites use numbers without explaining them, and that creates the illusion of objectivity where none exists.

An 83/100 looks measured. Without the formula and source data, it may only be an opinion wearing decimal-friendly clothes.

Our formula is published. Our weights are published. Our sources are tagged on every axis. If you think we got something wrong, you can check the math yourself.

That is the standard readers should hold us to.

---

*Maxwell Trent is a product researcher at EcomAgentTools. He designed the audience-level system and scoring rubric. He previously ran an ecommerce venture that didn't work out, which taught him more about tool evaluation than any success ever did. You can find all 83 tools with their scored breakdowns at ecomagenttools.com.*

More ecommerce AI guides