Trust Score Methodology
Our transparent 4-pillar scoring system helps you make informed decisions about AI tools.
Editorial Independence
Every score, badge, award, and verification result on Intelloro is produced under these binding editorial rules. No exceptions, no carve-outs for advertisers.
- Scores and ranking cannot be purchased — sponsorship never affects a tool's Trust, Task, or Dimension scores, nor its position in ranked or sorted results. Sponsored placements are always clearly labeled “Sponsored” and never receive inflated scores.
- Our scores are our own analytical opinion, not an independent audit — they are AI-assisted assessments, reviewed by a human, based on publicly available evidence. Compliance certifications are shown as claimed by the vendor and cross-checked against public sources — Intelloro does not itself audit, certify, or independently test any product, and the scores are not verified user ratings.
- Editorial awards cannot be bought — Editor's Pick + Featured selections are made through Intelloro's founder-led editorial review, based on verification coverage and merit, independent of sponsored placements.
- No marketing language is treated as evidence — compliance fields (HIPAA, SOC 2, MCP, etc.) require explicit certification statements with source URLs. "HIPAA-ready" and "SOC 2 in progress" never count.
- High-stakes facts are never guessed from AI memory — compliance certifications, customer names, funding, pricing, data residency, privacy, PII & data-deletion handling, security isolation, uptime, enterprise SSO, regional availability, copyright & licensing, and underlying-model attribution come only from the vendor's own site, live web verification, or our manual review. When no evidence exists the field is left blank — never filled from a language model's training knowledge (grounding-over-generation, aligned with the NIST AI RMF and ISO/IEC 8000 data-quality standards).
- No cross-tool contamination — data scraped from competitor comparisons or "Top 10 alternatives" articles is never used as evidence for the tool being evaluated.
- No third-party review ratings are publicly displayed — only npm and PyPI package-download counts are shown (npm public-API terms / PyPI CC BY 4.0). Review-platform star ratings are not reproduced.
- Thin content does not produce findings — pages with fewer than 300 characters of relevant content are marked UNVERIFIABLE, never silently filled.
How We Calculate Trust Scores
Every AI tool on Intelloro receives a Trust Score from 0 to 100 based on four weighted pillars. This score is designed to give you an objective, data-driven assessment of each tool's reliability, transparency, and user satisfaction.
Our methodology combines market-adoption signals (named notable customers), company and funding signals, operational-reliability signals (measured uptime, status page, data residency, incident record), independent verification, and self-reported compliance signals to create a comprehensive trust profile. Third-party review-platform star ratings (Capterra, GetApp, and others) are never fed into the Trust Score and are not publicly displayed anywhere; only npm and PyPI package-download counts are shown.
The 4 Trust Pillars
Each pillar contributes a specific weight to the total score of 100 points. An item is scored only on the signals that apply to it — a missing signal is never counted against it. Items with very little verifiable public data are marked "Insufficient data" (shown as "—") rather than given a misleadingly low score.
Compliance & Verified
Independent verification plus publicly claimed security, privacy, and regulatory compliance. "Verified" means the item passed Intelloro's verification pass; certifications reflect vendor statements (shown "as claimed"), not an independent Intelloro audit.
Operational
Evidence the product is reliable and transparent about how it runs. Uptime only earns points when backed by a public status page showing a measured percentage; a clean incident record earns credit, and a recent incident is penalized (time-decayed).
Market Proof
Evidence that real organizations rely on the product, via named notable customers. Third-party review-platform ratings are NOT used (their terms bar reproduction by a directory) and Intelloro has no first-party reviews yet, so this pillar is intentionally small; a self-reported customer list should not rival audited compliance. A first-party review signal will be added here when Intelloro's own reviews launch.
Company
Signals that a real, resourced company stands behind the product and is likely to keep it running.
Trust Tiers
Based on the total score, each tool is assigned a Trust Tier that gives you a quick visual indicator of overall trustworthiness.
Industry-leading trust and reliability. Top-tier in all categories.
Highly trustworthy with strong performance across most factors.
Reliable choice with solid fundamentals and room for improvement.
Acceptable but consider alternatives. Some trust factors need attention.
Significant improvements needed. Use with caution.
Not enough verifiable public data to score yet — shown as "—", never a misleadingly low number.
Dimension & Task Scores
Beyond Trust Scores, every tool and agent receives Dimension Scores and Task Scores, each rated 1–10. Tools are scored on 9 dimensions (6 universal + 3 AI-specific). Agents are scored on 12 dimensions (the same 9 plus 3 agent-specific dimensions). Task Scores are drawn from per-category rubrics described below.
The Scoring Dimensions (9 for Tools, 12 for Agents)
Grouped into three tiers: Universal SaaS (6 dims, all products), AI-Specific (3 dims, all AI products), and Agent-Only(3 dims, agents only).
How intuitive the tool is for new users. Considers onboarding flow, UI design, documentation quality, and learning curve.
The quality and reliability of results produced. Evaluates accuracy, consistency, and relevance of outputs.
Price-to-capability ratio compared to alternatives. Considers pricing tiers, feature access, and usage limits.
Flexibility to adapt the tool to specific needs. Includes API access, configuration options, and extensibility.
Quality of documentation, community resources, and customer support channels available to users.
How well the tool connects with other tools and workflows. Evaluates APIs, webhooks, and ecosystem compatibility.
Factual correctness, hallucination resistance, and consistency across repeated queries. AI-estimated from public product content; published benchmark scores are shown separately when the item is on a public leaderboard.
Compliance certifications (SOC 2, GDPR, HIPAA, EU AI Act), PII handling, copyright policy, and accessibility (WCAG). Aggregated from verified compliance signals.
Response time and inference speed. Derived algorithmically from the item's documented response-time tier.
Agent-only — Success rate at completing multi-step tasks autonomously. AI-estimated from public product content.
Agent-only — Accuracy when calling external tools, APIs, and functions. Critical for agents that use MCP, function-calling, or tool orchestration.
Agent-only — Sophistication of multi-step reasoning, goal decomposition, and dynamic replanning. NONE / BASIC / ADVANCED capability clamped.
Task Scores — Per-Category Rubrics
Task Scores rate how well a tool or agent performs at specific tasks within its category. Intelloro maintains a per-category set of scoring rubrics — one per vertical category. Each rubric defines 10 task capabilities that are the prime evaluation axes for that vertical (e.g., a Search & Discovery rubric scores search accuracy, semantic understanding, real-time indexing, filter refinement, source attribution,and 5 more). Each task is scored 1–10 against a 4-band anchor (1–3 / 4–6 / 7–8 / 9–10) — modeled after per-category feature-rating patterns used across the software-review industry.
Rubric selection is deterministic: an item's category_id maps one-to-one to its scoring rubric (e.g., data-analytics → the Data Analysis & BI rubric), so every item in a category is graded on the same 10 canonical tasks for fair, apples-to-apples comparison. Rubric definitions are locked in code and protected against drift by automated audit checks.
How Scores Are Generated (AI-Assisted)
Every tool and agent is scored by an AI-assisted pipeline that reads the vendor's own website (pricing, features, documentation, security pages), cross-references the data with the independent verification sources described above, and grades each dimension and task against the canonical 4-band rubric anchor. Scores are not vendor-supplied — they are independently generated from public evidence by our own model, reviewed by a human, and they update whenever the underlying data changes.
How Our Category Taxonomy Is Built
The rubrics are the fixed comparison targets, but the free-text task labels our extraction pipeline produces (and admins enter) vary in wording. We normalize them through a controlled alias map — a curated set of aliases — so synonyms like “content writing” and “copywriting” resolve to one canonical rubric, keeping every item comparable. Aliases are grounded in recognized, independent standards rather than vendor marketing: W3C WCAG 2.2 for accessibility tasks, the NIST AI Risk Management Framework for governance, OpenTelemetry semantic conventions for observability, and the DAIR.AI taxonomy for prompt engineering — supported by public reference encyclopedias for general software categories. Every alias must map to an existing rubric: a continuous-integration drift gate rejects any that points to a non-existent category, keeping the map and the rubric set in lockstep, and new aliases require a cited source before they are merged.
Why This Methodology
The 60-rubric architecture aligns with the per-software-category feature-rating patterns used across the software-review industry. Forrester Wave and Gartner Magic Quadrant use analyst-determined criteria per evaluation, which is more nuanced but does not scale beyond ~50 vendors per cycle. Intelloro's LLM-judged fixed-rubric approach is optimized for breadth — every item in every category is graded on the same canonical task set, keeping comparisons consistent across thousands of items.
Data Sources
Every data point on Intelloro is tagged with its source so you know where the information comes from.
Data verified through Intelloro's verification process — direct inspection of public vendor sources, cross-referenced.
Data automatically extracted from the product website using AI analysis.
Data cross-referenced using AI-powered web search for accuracy.
Technical data sourced directly from the GitHub API (stars, license, etc.).
Data analyzed from publicly available product information using AI.
Confidence Levels
Each score and data point is assigned a confidence level indicating how reliably it has been verified.
Data verified through multiple reliable sources. You can rely on this information with high confidence.
Data sourced from product website or documentation. Generally reliable but may need re-verification over time.
Data could not be fully verified from available sources. Treat as approximate and check the product website for the latest information.
Awards & Recognition
Intelloro recognizes exceptional AI tools and agents through editorial awards. These awards highlight products that stand out in quality, data completeness, and user value.
Editor's Pick
Our highest editorial recognition. Awarded to tools and agents that demonstrate excellence across multiple dimensions.
Featured
Highlighted on the homepage and category pages. Awarded to tools and agents that offer strong value in their category.
Trending
Automatically assigned based on recent visitor interest, search volume, and engagement metrics.
Editorial Discretion
Editor's Pick and Featured selections are made through Intelloro's founder-led editorial review. While we consider quantitative signals (Verification Coverage Score, verification status, dimension scores), final selections involve editorial judgment. Awards cannot be purchased and are independent of sponsored listings.
Verification Levels — What the Badge Means
Every listing carries a verification-level badge (Fully Verified, Largely Verified, Partially Verified, Limited Data, or Unverified) reflecting how many of its data points we have completed via our extraction + independent verification pass. The badge describes the state of our data about a tool, not a rating of the tool itself.
Fully Verified
≥ 90% of data points verified — comprehensive coverage across all tiers.
Largely Verified
70–89% verified — strong coverage with minor gaps in optional fields.
Partially Verified
50–69% verified — core information present, several optional fields incomplete.
Limited Data
30–49% verified — basic information only, significant gaps remain.
Unverified
< 30% verified — minimal data available; pending our extraction + verification pass.
For vendors: if your listing carries a lower verification level, you can request faster verification via the claim flow. The badge updates automatically after each extraction + verification pass. Industry pattern reference: LinkedIn Profile Strength (5 levels), Crunchbase Profile Strength, Wikipedia Content Assessment letter grades.
External Verification
Verified listings on Intelloro are cross-referenced against a dozen-plus independent external sources. Not every listing has completed this process yet — each listing shows a verification-level badge (see above) indicating how much of its data has been cross-checked. This cross-referencing brings accuracy beyond what vendor websites alone can provide.
Adoption
- npm — package download counts
- PyPI — package download counts
- Review-platform ratings (Capterra/GetApp/G2/etc.) are not stored or displayed.
Compliance
- SOC 2 certification verification
- GDPR compliance verification
- HIPAA compliance verification
- Status page & uptime verification
Adoption & Integrations
- Enterprise customer verification
- MCP compatibility check
- Make.com integrations
- npm / PyPI download counts
- Uptime SLA verification
- Zapier integrations
Data Quality & Vocabulary Maintenance
Every classification on Intelloro — task categories, supported countries, supported languages, data residency regions — runs through a controlled vocabulary and a multi-layer normalizer before reaching a tool or agent listing. The same rules apply to every entry on the platform: claimed, unclaimed, free, paid, sponsored or organic. Vendors cannot purchase a different vocabulary, override a normalizer rule, or skip the drift-detection cycle. The pipeline is summarized below.
Controlled vocabulary
Free-text task labels collapse to a curated set of canonical task-score rubrics. New aliases are added only after they recur across multiple tools and are backed by a cited standard.
- • 10 task keys per category rubric
- • CI-gated drift protection
- • Industry parallel: Algolia synonyms
Multi-layer normalizers
Country, language, data-residency, and task-category values flow through a progressive normalization chain before vocabulary lookup, so equivalent values written in different formats resolve to one canonical value.
- • Country — 200+ canonical ISO names
- • Languages — canonical language names
- • Data residency — canonical region tags
Drop-detection cycle
When a label or value can't map to the vocabulary, it's logged to a drift collection (not silently dropped). A scheduled job re-extracts canonical values from cached tool data and surfaces persistent misses for editorial review.
- • Runs on a regular schedule
- • Batched and cost-capped
- • Industry parallel: Schema.org freshness
Non-bias guarantee
Tools and agents are processed by the exact same pipeline. Cross-collection symmetry — the resolver, the vocabulary, the cron and the audit gates all treat the two entity types identically — means a vendor cannot get preferential classification on either side. Sponsorship, paid tiers, and editorial awards never override the rules below.
The same controlled vocabulary, normalizers, and audit gates apply to every listing — there is no paid pathway to bypass them.
AI Origin Classification — How We Label AI-Native vs AI-Layer
For some tools and agents, Intelloro publishes an AI Origin label. This is Intelloro's editorial opinion — a transparent, subjective classification based on publicly available information, not a statement of fact and not a claim that any vendor misrepresents its technology. We apply one simple, disclosed test.
The “Remove-the-AI” test
Imagine the product with its AI removed — what's left? Nothing usable, the AI is the product → AI-Native. A complete, standalone product remains → AI-Enhanced. Little but a thin shell over a third-party model → AI-Layer.
AI-Native
Built AI-first from day one; the AI is the core product, often with the vendor's own model or research (e.g. ChatGPT, Claude, Midjourney, Perplexity).
AI-Enhanced
An established product that existed before AI and added AI features later; remove the AI and a full product remains (e.g. Notion, Grammarly, Adobe, Salesforce).
AI-Layer
A thin layer built primarily on a third-party model, with limited proprietary technology. This is a neutral structural description, not a judgment of quality or value — many AI-Layer products are excellent and well-loved.
This label reflects Intelloro's opinion from public information and may be updated as products evolve. Vendors who believe a classification is inaccurate can request a review via our contact page.
Frequently Asked Questions
How often are Trust Scores updated?
Trust Scores are recalculated weekly (every Saturday) and immediately whenever a tool's underlying data changes — for example, after a new verification pass or a data update.
Can vendors influence their Trust Score?
Vendors cannot pay to improve their score. They can improve it by earning strong ratings on external review platforms, publishing a public status page with real measured uptime, maintaining a clean incident record, disclosing recognizable customers, funding, and team size, being transparent about data residency, completing Intelloro's verification pass, and obtaining compliance certifications (GDPR, SOC 2, ISO 27001/27701, HIPAA, FDA, WCAG, EU AI Act) — the same signals our 4-pillar model measures.
Why does a popular tool have a lower score?
Popularity doesn't guarantee trust. A tool might have many users but weak operational transparency (no public status page), privacy or compliance gaps, or little independent verification. Our score reflects the complete picture.
How does Intelloro verify a listing's data?
Every listing is first built by our automated extraction pipeline (which combines vendor-site content with live web-search grounding). Eligible listings then go through our independent verification pass — an AI-assisted review that re-checks the data against a dozen-plus independent public sources, with each session reviewed and approved by a human before it is applied. Every listing shows a verification-level badge reflecting how much of its data has completed this process.
What if I disagree with a Trust Score?
You can submit feedback through our contact form. If you have evidence of inaccurate data, we will investigate and update accordingly.
How are Editor's Pick and Featured awards decided?
Editor's Pick and Featured selections are made through Intelloro's founder-led editorial review, based on verification coverage, verification status, and overall product merit. While quantitative signals guide us, final selections involve editorial judgment. Awards cannot be purchased.
How can I display an Intelloro badge on my website?
If your tool has been recognized as Editor's Pick, Featured, or Verified on Intelloro, you can embed a badge on your website. Visit our badges page at intelloro.com/badges for embed codes and badge images.
How are Dimension Scores calculated?
Tools receive 9 Dimension Scores (6 universal — Ease of Use, Output Quality, Value for Money, Customization, Support, Integration — plus 3 AI-specific — Accuracy & Reliability, Compliance & Data Protection, Performance). Agents receive 12 Dimension Scores (the same 9 plus 3 agent-only: Task Completion, Tool Use Correctness, Planning Quality). Each is rated 1-10 using an AI-assisted pipeline that reads the vendor's own documentation and grades against canonical 4-band rubric anchors (1-3 / 4-6 / 7-8 / 9-10). Scores update when the underlying data changes.
How are Task Scores calculated?
Intelloro maintains 60 fixed scoring rubrics — one per vertical category — with exactly 10 canonical tasks per rubric (600 total task definitions). An item's category_id maps deterministically to its rubric (e.g., 'data-analytics' → the Data Analysis & BI rubric), so every tool in a category is graded on the same 10 canonical tasks for fair, apples-to-apples comparison. Scores are 1-10 with the same 4-band anchor used for Dimension Scores. The architecture mirrors industry per-category feature-rating models but scales via LLM-as-judge automation rather than analyst hours.
What do confidence levels mean?
Each data point on Intelloro has a confidence level — High, Medium, or Low — indicating how reliably it has been verified. High confidence means data was confirmed through multiple sources or direct verification. Medium means it was sourced from the product website. Low means it could not be fully verified from available sources.
Have Questions?
If you have questions about our methodology or want to report an issue with a Trust Score, we'd love to hear from you.
Contact Us