Marketplaces & Two-Sided Platforms

Your platform publishes other people's links at industrial volume

A marketplace is a distribution channel that anybody can write to. Listings, storefronts, profile bios and buyer-seller messages all carry hostnames chosen by strangers, and every one of them arrives with your brand around it. This page covers where a phishing check belongs in that pipeline, and why the arithmetic pushes you toward a local feed rather than a call per listing.

390,000+Live phishing domains, every one DNS-verified
100 / POSTDomains per batch call, one lookup each
10 req/sRate limit per API key — the number that shapes the design
Sub-50 msResponse time, so a submission check stays inline

Where a hostname enters

Listing bodySpec sheets, manuals, size guides, "see our other stock" links.
Storefront profileSeller bio, returns policy page, social handles, external shop.
Buyer-seller threadTracking links, invoice links, "pay here instead" links.
Mail impersonating youListing-reported notices and payout-verification prompts.
Listing submission Trust & safety queue Payout settings Appeals Seller API partners Search & category pages
Home / Use Cases / Online Marketplaces
The intermediary position

You are liable for a page you did not write

A single-brand retailer controls every link it ships. A marketplace controls the frame around links that arrive from an unbounded population of strangers, at a rate no review team can match.

The structural problem is not that some sellers are dishonest. It is that a marketplace's core product is a write path into a trusted surface. Anyone who completes registration can publish text, images and hyperlinks that appear inside your chrome, under your typography, next to your buyer-protection badge, and ranked by your search. Every design decision that makes selling frictionless — instant listing, bulk upload via a partner integration, editable descriptions, rich seller storefronts — widens that write path. The friction you remove for honest sellers is removed identically for the other kind.

A liability shape frameworks miss

When a buyer types card details into a look-alike checkout reached from a listing on your site, the buyer's account was not compromised, your infrastructure was not breached, and no control you own failed in a technical sense. The buyer still experienced the harm on your platform, will describe it that way in a review and to a regulator, and will expect a refund. Trust and safety absorbs the case, payments absorbs the chargeback exposure, and the brand absorbs the story — none of it affected by how well your own perimeter is defended.

The fraud runs in both directions

Buyer-side attacks harvest saved payment methods, addresses and order histories. Seller-side attacks are frequently more valuable per compromise, because an account with trading history, positive feedback and verified status is an asset that took months to build and can be monetised immediately — by listing goods that will never ship, by changing the bank account that receives settlement, or simply by selling the account on. A platform that only defends buyers is protecting the cheaper half of its own exposure.

The impersonation channel you do not own

Your operational emails are the most predictable, highest-trust messages your users receive: policy update, listing removed, item reported, verify your payout details, action required within 48 hours. Those templates are public — every seller has received dozens — and reproducing them costs nothing. The resulting message lands in a seller's inbox with your logo and a hostname that is not yours, at a moment when the seller genuinely does have listings that could plausibly have been reported.

Make the hostname cheap to evaluate

The practical response is not to review everything. It is to make hostnames a fixed, cheap thing to check wherever they appear: at listing submission, at message send, at profile save, at payout-change, and in the moderation queue where a human is already looking. A known-bad lookup against a DNS-verified list of live phishing hosts does not tell you whether a listing is fraudulent — it tells you in single-digit milliseconds whether the destination is already confirmed hostile, which is a small answer that removes a large amount of queue volume.

Listing descriptions Image alt text and captions Reviews and Q&A Seller bios Shipping and tracking notes Community forum posts All of it user-controlled
Six recurring shapes

What actually shows up in a marketplace abuse queue

Each of these has a different owner internally, which is exactly why they get defended inconsistently. Domain checking is the one control that applies to all six identically.

Seller side

Storefront account takeover

An established seller with trading history is worth far more than a new one, so the lure is aimed at the fear of losing it: a suspension notice, a policy violation, an appeal window that closes tomorrow. Once inside, the attacker edits payout details, bulk-relists high-value goods at plausible prices, and lets the account's own reputation do the persuading.

Buyer side

Buyer credential harvesting

A cloned sign-in page collects the address book, the saved cards and the order history in one step. The order history is the underrated part: it tells the next message exactly what the victim bought, from whom, and when, which turns a generic follow-up into a specific and convincing one about a real purchase.

Payments

Off-platform checkout lures

The seller offers a discount for paying directly, or claims the platform's payment step is broken, and supplies a link to a checkout that mimics yours closely enough to pass a glance. The buyer loses protection, the dispute is unwinnable, and your support team fields a case about a transaction that never touched your ledger.

Content

Links buried in listing bodies

A specification PDF, a sizing chart, a "full catalogue" link, a warranty registration page. Long-tail categories with technical products give a seller a legitimate-sounding reason to link outward, and the link frequently appears in the third paragraph of a listing nobody reads until something goes wrong.

Messaging

Thread-context abuse

The message system is trusted because it is internal and because the buyer is expecting to hear about their order. A tracking link, a delivery reschedule, a customs fee — each is a plausible reason to leave the platform mid-transaction, and each arrives inside a conversation the buyer themselves started.

Impersonation

Payout and "reported listing" mail

Two templates carry most of the seller-side volume: a payout that cannot be released until banking details are re-verified, and a notification that a listing has been reported and requires acknowledgement. Both create urgency without asking for money, which is what gets them past the instinct that catches cruder attempts.

Reading down those six, the common factor is that every one resolves to a hostname the user is about to visit. That is a much narrower question than "is this seller fraudulent", and narrow questions are the ones that can be answered automatically at volume. It is also why the same check can serve teams that otherwise share nothing: content moderation, messaging abuse, payments risk and the corporate mail gateway are four different backlogs, four different tools and usually four different reporting lines, but they are all asking whether one string is a known hostile host.

Off-platform

The lure that moves the payment, not the account

The most damaging marketplace attack does not steal a password at all. It persuades a buyer that the safest next step is to pay somewhere else.

Off-platform payment fraud is uncomfortable to defend against because the manipulation is entirely social and the technology involved is trivial. A seller opens a conversation about a genuine listing, establishes normal rapport, then introduces a reason to complete elsewhere: a fee that can be avoided, a stock system that is not syncing, an invoice that must be settled directly for a business purchase. The buyer receives a link to a page that carries your colour scheme, your button shapes and often your literal stylesheet, because copying the front end of a checkout is a matter of saving a page.

The buyer is being cooperative, not careless

In their own frame of reference they are helping a seller they have been talking to for two days about an item they want. The signals that normally trigger suspicion — an unexpected message, a stranger, a request out of the blue — are all absent by design. Awareness banners underneath the message box help at the margins and are worth having, but they compete against a conversation that has already established trust, and they are read least by exactly the buyers most likely to comply.

Intervene at send, not at click

When a message containing a hostname is submitted, the platform already has the string in hand and already runs it through spam heuristics. Adding a known-bad domain lookup there costs a single-digit-millisecond decision and yields a hard, non-probabilistic answer: this host is confirmed live phishing infrastructure, or it is not on the list. A confirmed match is not a case for a moderator — it is a message that never delivers, an automatic account flag, and a signal worth propagating to that seller's other listings and threads.

Write the interstitial like it matters

If a buyer does click outward, the warning page is the last control you own and its wording decides whether it works. "You are leaving our site" is a disclaimer that trains people to click through. A statement that this specific address has been verified as an active fake payment page, that no purchase protection applies once they leave, and that your own checkout is the only place the order can complete, is a different piece of writing — and its accuracy is what earns attention, which is why it should appear only for confirmed matches.

Platforms that want to reason about the shortener layer, which is heavily used in exactly this pattern, will find the specifics on the URL shorteners page; the sending side of the impersonation problem is covered under email security.

"Pay direct and skip the fee" "Our checkout is down today" "Business invoice, pay by transfer" "Customs fee before delivery" "Reschedule your delivery here"
Where the check sits

Four points in the submission path, all of them server-side

The check belongs where you already hold the string and already have a decision to make, not in a browser and never in code the seller can read.

Extract at parse time

Pull every hostname out of the listing body, the storefront fields and any attachment metadata as part of the same pass that already sanitises markup. Deduplicate before anything else — one listing rarely holds more than a handful of distinct hosts.

Look up locally first

Match against the copy of the database you already hold from the daily feed. The overwhelming majority of submissions resolve here with no external call at all, which is what keeps the submission path fast and the credit spend flat.

Batch what is left

Hosts not seen since your last pull go into a queue and out in batches of up to a hundred per request, one lookup each. The response returns checked, phishing_found and credits_used, which is enough for both the decision and the audit line.

Act, then propagate

Block the submission, hold it for review, or publish with the link stripped. Then apply the same verdict to the seller's other listings, open threads and storefront fields rather than treating each occurrence as an isolated event.

Detail one: the key never ships to a client

The API key is the username chosen at registration, which makes it a server-side secret with no exception: it must never be embedded in a mobile app bundle, a single-page application, a seller-facing widget or anything else that reaches a browser. If you want sellers to be able to check a link themselves, put a thin endpoint of your own in front of it, rate-limited per account, and keep the credential where only your servers can see it.

Detail two: deduplicate or overpay

Marketplace content is enormously repetitive — the same manufacturer specification page appears across thousands of listings in a category, the same courier tracking host appears in nearly every shipping message, the same handful of social platforms appear in most seller bios. A cache keyed on hostname with a lifetime tied to your feed refresh collapses that traffic to almost nothing. Without it you will pay repeatedly to be told the same thing about the same host, and burn rate limit doing it.

The arithmetic

Ten requests per second is the number that shapes the design

Work the volume through honestly and the architecture decides itself, well before anybody argues about price.

10 / sRequests per API key, per second
1,000 / sDomains per second when every request is a full batch of 100
86.4 MDomains per day at that sustained ceiling
390,000+Entries in the database you can simply hold locally

Why per-listing API calls stop scaling

A platform taking a hundred thousand listings a day with an average of two distinct external hosts each needs two hundred thousand lookups before any messaging traffic is counted. Sent one at a time, that is a sustained request rate you cannot reach on a single key and a credit line that grows linearly with your own success. Batched at a hundred per call it is two thousand requests — trivially inside the limit, but still billed per domain, and still an external dependency sitting inside your publish path.

Why a local copy inverts the problem

Download the full database once a day and the same two hundred thousand lookups become in-process comparisons against a set that fits comfortably in memory. Feed downloads are unlimited with a subscription, so the cost of a busy day is identical to the cost of a quiet one. The API then handles only the residue — hostnames genuinely unseen since your last pull — which is a small fraction of a repetitive corpus and sits far inside the rate limit even during a listing surge.

Resilience beats cost as an argument

An inline external call means a dependency whose latency and availability now sit inside your seller experience. When that service is slow, listings are slow. When it is unreachable, you have to decide in advance whether submissions fail closed and sellers cannot list, or fail open and the check silently stops happening — and whichever you choose, you discover the consequences at the worst possible moment. Matching against a local file removes the choice: local operation, local failure modes, and a missed pull degrades freshness rather than availability.

The daily rhythm to build against

The database is rebuilt every 24 hours and the build lands at 04:30 UTC. A scheduled job pulls the current CSV — domain,category,dns_status — or the JSON equivalent, applies the changelog of domains added and removed rather than reimporting everything, swaps the in-memory set atomically, and emits a metric. One alert if the file is older than forty-eight hours is the only monitoring this needs, and it catches the single realistic failure: a job that stopped running while everything downstream carried on reporting success.

Trust and safety

Queue enrichment, and what an appeal reviewer needs

The value here is not only the automatic block. It is the field that stops a human spending nine minutes establishing something a lookup answers instantly.

Any moderation queue that runs on human judgement is a throughput problem before it is an accuracy problem. Reviewers work under a handle time, decisions are sampled for quality, and the expensive cases are not the obvious ones — they are the ambiguous items where the reviewer opens a second tab, looks at a domain, forms an impression, and makes a call they could not fully justify afterwards. Multiply that by a queue and the cost is enormous, and the consistency between reviewers is worse than anybody wants to admit.

A fact beats an impression

When a queued item carries a field stating that the linked host is confirmed active phishing infrastructure, with a last_checked date and an explicit confidence of 0.98, the reviewer is not forming an impression; they are applying a policy to a fact. Handle time on those cases collapses, the decision is reproducible across reviewers, and the enforcement is far easier to defend later. Items whose hosts are simply absent should be labelled "not on the list" rather than "clean", because a reviewer who reads clean will stop looking.

Appeals are where evidence decays

A suspended seller writes in and says the link was to a legitimate supplier page. Without a stored verdict, the appeals reviewer is reconstructing history from a domain that may have been taken down, repurposed, or may resolve perfectly well now — and in that situation the sympathetic answer is to reinstate. With a dated verdict captured at decision time, the appeal is answerable in one line, the sequence is auditable, and reinstatements happen where they should rather than where the record has faded.

Design the innocent case too

Sellers do get caught by a domain that was compromised while they were legitimately linking to it — a small supplier's site taken over is the classic instance, and the seller genuinely did nothing wrong. The right handling is to remove the link and notify rather than suspend the account, and to re-check on a schedule so the restriction lifts on its own. Because entries are DNS-verified and fall out when they stop resolving, that recovery happens naturally instead of requiring a manual amnesty.

Platforms feeding this into a wider detection stack usually land it in the same pipeline described under SIEM integration, and the takedown side of brand impersonation is covered under brand protection.

Choosing the integration

Per-lookup calls, a local feed, or both

Most platforms end up with both, but for different traffic. The point of the table is to make the split deliberate rather than accidental.

Decision factorPer-lookup API callsDaily feed held locallyBoth, split by traffic
Latency inside the publish pathSub-50 ms, but a network hop you now own In-process comparison, no hop Local for the common case
Cost as listing volume grows Scales linearly with your success Flat — downloads are unlimited Flat plus a small residue
Behaviour when upstream is unreachable You must pre-decide fail-open or fail-closed Unaffected; freshness degrades, not availability Falls back to the local set
Freshness between daily builds Always the current database As current as your last pull Current where it matters most
What leaves your networkOne hostname per call Nothing about user activityOnly unmatched residue hostnames
Rate-limit headroom during a surge Tight — 10 req/s per key is the ceiling Not applicable to local matching Comfortable; residue is small
Best fitOnboarding checks, appeals, ad-hoc investigationListing, messaging and profile submission at volumeThe realistic answer for a live platform
Feed downloads are unlimited with a subscription; API calls consume one credit per domain checked.

The split most platforms converge on is simple to state. High-volume, repetitive, latency-sensitive traffic — every listing body, every message, every profile save — matches against the local set. Low-volume, high-consequence checks that must reflect the database as it is right now — a seller onboarding decision, an appeal review, an investigation into a suspected ring — go through the API, because a handful of extra credits is irrelevant against the cost of being wrong on a case a human is already reading.

State the limitation as precisely as the architecture

A platform that misunderstands this will build the wrong policy on top of it. This is a known-bad lookup: a confirmed match returns a confidence of 0.98 and a category of phishing/malware; anything else returns 0.0. There is no intermediate score, no verdict on pages the database has not seen, and no coverage for a domain registered and first used inside the same hour. Treat a non-match as "not on the list", route it back to your existing heuristics and reputation signals, and never let it short-circuit a review that would otherwise have happened.

What it costs

Sizing the spend against a platform's actual traffic

Two mechanisms, priced differently on purpose, and the sensible allocation is not the one most teams reach for first.

Credits are monthly subscriptions via PayPal and billed monthly, with unused credits expiring at that point. The published tiers run from Growth at $99/month for 25,000 lookups, through Growth at $99/month for 25,000 lookups at $0.0040 and Professional at $249/month for 100,000 lookups with priority support, up to Business at $499/month for 250,000 lookups with dedicated support, Enterprise at $999/month for 750,000 lookupsand Enterprise at $999/month for 750,000 lookups at $0.0010, both with an account manager, and Enterprise at $999/month for 750,000 lookups on the Enterprise plan with an SLA. A fourteen-day refund applies where under ten percent of a package has been used — which is the relevant path for most platform-scale purchases, since finance departments at this size rarely want a card transaction of that size.

Why the allocation is obvious

A platform processing two hundred thousand distinct listing hostnames a day would exhaust the largest monthly plan in roughly a month if every one went through the API. The same platform holding the database locally pays $499 per month for the Daily Threat Feed, with unlimited downloads and no relationship between traffic and bill. The annual option saves $1,989 and adds archive access, priority support, custom formats, an account manager and 100,000 credits — which is, in practice, exactly the budget for onboarding, appeal and investigation traffic.

When delivery has its own requirements

The enterprise route covers SFTP or S3 delivery into an existing ingestion pipeline, STIX/TAXII for a security operations centre that already consumes structured threat data, a custom update frequency, an SLA, and on-premise deployment. Those are quoted rather than listed. The current credit table is on the pricing page and the delivery formats are set out under daily feed.

A health check that costs nothing

The /stats endpoint reports database size and last update time without an API key, without credits and without authentication. That makes it a legitimate external probe for your ingestion pipeline — a synthetic monitor can confirm the upstream build advanced today, independently of whether your own pull succeeded, which distinguishes "the source is stale" from "our job is broken" without consuming anything.

Questions from platform teams

What engineering, trust and safety, and risk ask

Including the objections that do not resolve neatly, since those are the ones that determine whether this survives a design review.

Can we call this inline on every listing submission?

Technically yes — responses come back in under 50 ms — but at platform volume you should not want to. Every inline external call is a dependency inside your publish path, and 10 requests per second per API key is a real ceiling once a bulk-upload partner pushes a few thousand listings in a burst. The pattern that holds is local-first: match against your copy of the daily feed in process, and send only the hostnames you have not seen since the last pull to the batch endpoint, a hundred per request. That keeps the submission path free of network dependencies and keeps the API traffic small enough that the rate limit never becomes a design constraint.

Will this catch a phishing page hosted on a legitimate platform we cannot block?

Only where the hostname itself is in the database, which for a page hosted under a large shared domain generally means it is not. This is a hostname-level control, so a credential-harvesting page sitting on a mainstream file-sharing, form-building or site-builder service under that service's own domain will not be blockable without collateral damage you would never accept. That is a genuine gap and it is worth planning for explicitly rather than discovering it. The layers that handle it are different — content inspection, sender reputation, behavioural signals on the seller account, and the abuse-reporting relationships you hold with those platforms. Domain checking removes the dedicated hostile hosts, which is a large share of the volume but expressly not all of it.

What do we do about link shorteners in messages?

Resolve them yourself before checking, on your own infrastructure, with a redirect cap and a timeout, and then check the final hostname. Checking the shortener's own domain tells you nothing useful, since the same host serves millions of legitimate links. Many marketplaces go further and simply disallow shortened links in buyer-seller threads, because there is almost no honest reason to obscure a destination inside a conversation about an order. If you do allow them, expand and display the real hostname to the recipient rather than the shortened form. The mechanics are covered in more depth on the URL shorteners page.

How do we avoid punishing a seller whose supplier site got compromised?

Separate the link decision from the account decision, and make that separation explicit in policy. Remove or disable the link, notify the seller with the specific hostname and the date it was verified, and leave the account alone unless there is a second signal — repetition after notice, a payout change in the same window, a pattern across multiple sellers. Then re-check on a schedule. Because every entry is DNS-verified and falls out when it stops resolving, a compromised host that gets cleaned up disappears from the database on its own, and your restriction can lift automatically rather than waiting for somebody to process an appeal.

Does this help with fake versions of our own sign-in page?

It gives you verification rather than prevention, and the distinction matters. You cannot block a page for a buyer who opens it on their home broadband. What you can do is confirm, with a dated verdict, that a candidate hostname is currently live phishing infrastructure — which is the evidence a registrar or hosting provider wants before acting, and the basis on which you decide whether to warn users. Inside your own estate the block does apply: staff machines, contractor access, seller-support terminals and anything else resolving through your infrastructure. Combining the two is the approach described on the brand protection page.

We already run a commercial reputation service. Why add this?

Because they answer different questions and fail differently. Reputation scoring produces a probability from many signals, which is powerful and inherently fuzzy — it needs tuning, it drifts, and a reviewer has to interpret it. This produces a binary verdict against hosts confirmed to be resolving right now, with no interpretation required. In a moderation queue those complement each other well: the binary verdict resolves the cases it covers instantly and cheaply, and your reputation signals do the work on everything else. If your evaluation criterion is which single product covers the most ground, this is not that product and does not claim to be.

How stale can the local copy get before it matters?

The build cadence is 24 hours with the daily build landing at 04:30 UTC, so a once-daily pull is aligned with how fast the data actually changes and pulling more often mainly buys reassurance. Since downloads are unlimited under a subscription, pulling twice daily costs nothing if it makes your risk team more comfortable. What genuinely matters is knowing when the pull stopped. Track the age of the loaded file as a first-class metric, alert past forty-eight hours, and use the keyless /stats endpoint as an independent check on the upstream so you can distinguish a stale source from a broken job of your own.

Can we expose a link checker to our sellers and buyers?

Yes, provided the call happens on your servers. The API key is the username chosen at registration, which makes it a credential rather than a public identifier, and it must never appear in a mobile bundle, front-end JavaScript or page source. Put your own endpoint in front of it, rate-limit per account, and log the queries. Be careful with the wording of the result. "Not found in the database" is accurate; "this link is safe" is not, and a user-facing tool that implies the latter will eventually be quoted back at you in a dispute over a page the database had never seen.

Put the check where the hostname enters, not where the user clicks

Hold the database locally for listing, messaging and profile traffic, and keep credits for onboarding decisions, appeals and investigations. The feed page covers delivery formats; the API documentation covers the batch contract and the response fields.