Skip to main content

How to Correlate Malware Domains at Scale

A phishing alert containing one disposable-looking domain is rarely a one-domain problem. The operator may have registered dozens of related names, moved infrastructure between providers, or staged the next wave under domains that have not yet appeared in telemetry. Knowing how to correlate malware domains turns an isolated indicator into an infrastructure investigation.

The goal is not to prove that every similar domain belongs to the same actor. It is to build defensible clusters from independent signals, assign confidence, and surface the next domains worth blocking, monitoring, or investigating. That distinction matters. Aggressive correlation creates noisy clusters and false positives. Conservative correlation can leave active infrastructure undiscovered.

How to Correlate Malware Domains: Start With a Seed

A seed can be a confirmed phishing domain, command-and-control hostname, malware configuration artifact, suspicious redirector, or domain named in an incident report. Before pivoting, capture the seed's state at the time it was observed. Domain infrastructure changes quickly, and current DNS alone can misrepresent historical relationships.

Record the fully qualified domain name, registrable domain, first-seen and last-seen times, resolution history, current and historical IP addresses, nameservers, registrar, registration and expiration dates, certificate details, HTTP behavior, and passive DNS observations. If the domain was involved in phishing, preserve the URL path, page title, redirect chain, favicon hash, and targeted brand.

Normalize this evidence before correlating it. Lowercase hostnames, convert internationalized domains consistently, separate subdomains from registrable domains, and retain timestamps with every observation. A normalized dataset prevents the same infrastructure from appearing as several unrelated entities because one feed uses a trailing dot, another stores a punycode value, and a third reports only the apex domain.

Build Relationships From Multiple Signal Types

No single signal is sufficient for most production decisions. Shared IP space may indicate common hosting. Similar lexical patterns may simply reflect an industry trend. A useful cluster emerges when several signals agree within a meaningful time window.

Registration and naming patterns

Registration metadata can expose campaign construction. Look for domains registered within a narrow period, using the same registrar, nameserver pair, registration duration, or registrant pattern where that data is available and reliable. Newly registered domains that mimic the same brand, reuse uncommon terms, or follow the same token structure can be strong leads.

Lexical similarity should be treated as candidate generation, not attribution. For example, microsoft-security-check[.]com and microsoft-account-alert[.]com may be related, but the naming pattern alone does not establish common control. Confidence increases if both were registered on the same day, delegated to the same nameservers, and served matching phishing kits.

Modern registration privacy and inconsistent Whois coverage reduce the value of registrant fields. That is why normalized registration timelines, zone-level coverage, and first-seen timestamps are often more operationally useful than relying on an email address that may be redacted or fabricated.

DNS and hosting infrastructure

Passive DNS is usually the most productive correlation layer because it preserves historical associations. Pivot from the seed through A, AAAA, CNAME, MX, NS, and TXT records. Then examine whether related domains shared infrastructure concurrently or only years apart.

Shared nameservers can be particularly valuable when they are uncommon or campaign-specific. Shared IPs require more caution. A single virtual private server, bulletproof host, or dedicated reverse proxy can be highly discriminating. A large cloud provider, CDN, or shared hosting address is not. Weight the signal according to the infrastructure's exclusivity and the duration of overlap.

CNAME chains often reveal the operational layer behind a disposable hostname. Malware operators may rotate front-end domains while retaining a stable redirector, traffic distribution system, or hosting endpoint. Correlating on that stable middle layer can identify domains that look unrelated at the lexical level.

TLS and web artifacts

Certificate transparency data can connect domains through shared certificates, issuer choices, certificate request timing, subject alternative names, and repeated organizational fields. A wildcard certificate used across a small set of suspicious domains is a stronger signal than a common free certificate issuer.

Web artifacts add behavioral evidence. Compare page titles, favicon hashes, screenshot similarity, server headers, redirect destinations, JavaScript filenames, form action endpoints, and phishing-kit resources. Operators often change domains faster than they change templates. Reused assets can connect a credential-harvesting campaign even after it migrates to new hosting.

This evidence is time-sensitive. A domain may serve a phishing kit for six hours and then return a parking page. Store observations as events rather than overwriting them with the latest crawl result.

Malware and communication behavior

If the seed appears in malware telemetry, compare DNS request patterns, URI structures, User-Agent strings, TLS fingerprints, beacon intervals, and resolved endpoint behavior. Domains used by the same malware family are not automatically controlled by the same actor, but shared operational characteristics can improve a cluster's confidence.

Behavioral evidence is especially useful for separating infrastructure reuse from active coordination. Two domains on the same IP may be unrelated. Two domains that resolve to the same IP, use the same unusual URI path, and receive matching beacon traffic within the same campaign window are much more likely to be operationally connected.

Score Evidence Instead of Treating It as Binary

Correlation works better as a scoring problem than a yes-or-no rule. Give higher weight to rare, stable, and time-aligned relationships. Lower weight should go to common services, broad providers, and weak similarity measures.

A practical model may score a same-day registration window, shared uncommon nameservers, overlapping passive DNS records, identical phishing-kit assets, and a common redirect endpoint. It should also apply penalties for shared cloud or CDN infrastructure, stale associations, and relationships observed only once.

The exact weights depend on your threat model. For brand abuse monitoring, registration timing, lexical features, MX setup, and web content may matter most. For command-and-control tracking, passive DNS history, TLS fingerprints, and network behavior may be more valuable. Keep the score explainable so an analyst can see why two domains were clustered and reject weak edges quickly.

Use Time as a First-Class Correlation Field

A domain's associations are not permanent. Nameservers change, IPs are reassigned, certificates expire, and attacker infrastructure gets recycled. Without temporal context, correlation graphs become collections of misleading historical coincidences.

Require overlap where possible. If a malicious domain and a candidate domain resolved to the same dedicated IP during the same 48-hour window, that is meaningful. If they touched the same shared IP eighteen months apart, it may be irrelevant. Keep first-seen, last-seen, and observation timestamps on every relationship, then decay confidence as evidence ages.

Time also helps identify campaign waves. A burst of new registrations, followed by DNS activation, certificate issuance, and a cluster of phishing submissions, often indicates an operation before it reaches broad detection coverage. That sequence is more actionable than any single data point.

Operationalize the Workflow

The output of correlation should feed a repeatable security workflow, not remain a manually maintained investigation graph. Start by ingesting confirmed seed domains into a normalized domain intelligence layer. Enrich each seed with current and historical DNS, registration, zone, certificate, and web observations. Generate candidate domains through shared entities and similarity rules, then score and deduplicate the resulting relationships.

Send only high-confidence candidates directly to blocking, takedown, or incident-response queues. Route medium-confidence candidates into monitoring, sandboxing, or analyst review. Keep low-confidence candidates as graph context rather than turning them into alerts. This tiering prevents a useful discovery process from becoming an alert-volume problem.

For teams building this at scale, freshness and coverage are as important as correlation logic. Delayed zone ingestion misses newly registered domains during the period when they are most useful. Fragmented Whois sources create inconsistent joins. Brittle scrapers make a pipeline hard to trust during an active incident. A detection-ready domain dataset, such as the infrastructure provided by Primitive Host, reduces those collection and normalization gaps so correlation logic can run against consistent records.

Validate Clusters Before Taking Action

Correlation is an investigative accelerator, not a substitute for verification. Review the highest-impact clusters for false-positive risk before blocking a broad set of domains or attributing activity to an actor. Inspect whether the shared infrastructure is exclusive, whether the timing aligns, and whether web or malware behavior independently supports the relationship.

This is particularly critical around popular hosting platforms, free DNS services, certificate authorities, and URL shorteners. Those entities create dense graphs that can look suspicious simply because many unrelated domains pass through them. Exclude or downweight common nodes unless another stronger signal connects the domains.

A useful analyst view should show the evidence path, not just the final score: seed domain, shared nameserver, registration window, historical IP overlap, and matching page artifact. Analysts can make fast, defensible decisions when the system shows why a relationship exists.

The most valuable outcome is not a larger list of suspicious domains. It is earlier visibility into the small, coherent infrastructure set an operator is likely to use next. Build correlations that retain evidence, preserve time, and tolerate uncertainty. That gives responders room to act before the next domain becomes the next incident.

← Back to blog