Fail2ban has been the default answer to "my SSH box keeps getting hammered" for over a decade. It works, it is simple, and it knows nothing about the wider world — every instance is an island that re-learns the same botnet attacks from scratch. CrowdSec starts from the opposite premise: your server is one node in a global mesh, and when an IP attacks you, it should be blocked for everyone, and when it attacks everyone, it should be blocked for you before the first packet lands. That is "participative security," and it is the entire reason this 14,000-star Go project exists.
CrowdSec is a behavior detection engine plus an IP-reputation network plus a remediation layer, shipped as one binary under the MIT license. It reads your logs, parses them with Grok-style patterns, applies detection scenarios (if you see 5 failed SSH logins in 2 minutes from one IP, ban it), and then enforces a decision through a "bouncer" — a firewall rule, an Nginx 403, a Captcha, whatever fits. So far that is fail2ban with better engineering. The network is the twist: by default, the attacker IPs you detect are shared with the CrowdSec Network, and in return you receive a continuously updated blocklist of IPs the community has already flagged. You get smarter because the herd gets attacked.
This review explains what that architecture actually means to run, where the community network is a feature and where it is a caveat, how the WAF layer changes the game for self-hosters, and the privacy question that "open source security tool" write-ups tend to hand-wave: when you participate, what exactly leaves your box? By the end you should know whether CrowdSec belongs in your stack or whether you would rather keep your island.
1. The Architecture: Agent, CLI, Hub, and Bouncers
CrowdSec is not one program; it is a small system. The core daemon,
crowdsec, does the detection. It tails or receives logs (from files, from journald, from a syslog feed), runs them through parsers that turn raw lines into structured events, then evaluates scenarios — YAML-defined rules that say "when this pattern happens this many times in this window, emit this decision." Decisions are stored in a local database (SQLite for small installs, MySQL for larger) and exposed over a local API called LAPI.
cscli is the command-line control plane. You use it to see active bans, enable or disable parsers and scenarios, inspect the reputation of an IP, and manage your participation in the network. It is how you actually operate the system day to day — cscli decisions list shows you who is currently banned and why, which is oddly satisfying and genuinely useful for incident response.
The "Hub" is the distributed library of parsers and scenarios. CrowdSec ships collections for nginx, SSH, Traefik, WordPress, Postfix, and dozens more, maintained in a separate repository and pulled on demand. You enable the collections for the services you run, and CrowdSec starts understanding their logs. This modular design is why adding a new protected service is "enable a collection," not "write a regex from scratch."
The "bouncer" is the enforcement half. CrowdSec decides; the bouncer acts. There are bouncers for iptables/nftables firewalls, for Nginx and Traefik (injecting 403s or Captchas at the proxy), for the Linux firewall, for Windows, for Cloudflare, and more. The decision flows from the agent's LAPI to the bouncer, which installs the block. Decoupling detection from remediation is the key architectural choice: you can run one CrowdSec agent and have it feed bans to five different enforcement points.
2. How Detection Actually Works
The detection pipeline is worth understanding because it explains both CrowdSec's power and its failure modes. A log line enters a parser, which extracts fields — an IP, a status code, a user agent, a timestamp — into a structured event. A scenario then expresses a condition over those events:
evt.Parsed.status == "401" && count > 10 in 1 minute from same IP. When the condition trips, the agent emits a decision: ban that IP for a duration.
The scenarios are human-readable YAML, which is a genuine improvement over fail2ban's often-cryptic filters. You can read what a scenario does, tune its thresholds, or write your own in an afternoon. The Hub's community scenarios are battle-tested against real attack traffic, so out of the box you get baseline protection against the scan-and-probe campaigns that hit every public IP without any tuning. The project's own framing — "baseline detection is effective out-of-the-box, no fine-tuning required" — holds up for the common cases.
Where this bites is novelty. CrowdSec only detects what a parser understands and a scenario describes. If an attacker exploits a pattern your enabled collections do not cover, CrowdSec will not see it as an attack unless you write the scenario. The system is as good as the Hub content you have enabled plus your own custom rules. For mainstream services this is a non-issue; for a bespoke app with weird logs, you are writing parsers. That is the honest boundary of "works out of the box."
3. The CrowdSec Network: The Feature and the Caveat
Here is the part most reviews flatten. By default, when CrowdSec detects an attack on your machine, it shares the offending IP with the CrowdSec Network — a global, shared CTI (cyber threat intelligence) database. In return, your agent downloads a constantly refreshed list of IPs the community has flagged, and those get auto-banned at your edge. This is the participative model, and it is genuinely effective: a scanning campaign that hits thousands of nodes simultaneously gets blocked everywhere at once.
The caveat is data flow, and an operator who cares about sovereignty should name it precisely. When you participate, you send attacker IPs (and the attack type) to CrowdSec's infrastructure. You are not sending your logs, your users, or your content — only the hostile source IPs and the signal that they attacked you. That is a narrow, low-sensitivity stream. But it is an outbound connection, and "I share attack telemetry with a vendor" is a different posture from "my security tool is fully local." If your threat model includes "I must not emit any signal that my infrastructure exists or is under attack," you can disable network participation and run CrowdSec in standalone mode — you lose the community blocklist but keep local detection.
There is also a softer caveat: a shared blocklist is only as trustworthy as the community. False positives — a legitimate IP wrongly flagged by enough nodes — could get blocked at your edge. In practice the signal-to-noise is good and the project has abuse controls, but the principle stands: you are delegating part of your trust to a crowd. For most self-hosters that is a fine trade; for a high-security deployment, you tune participation and lean on your own scenarios. The point is that the trade is configurable, not forced.
4. The WAF Layer: AppSec at the Edge
CrowdSec's AppSec (WAF) component is what separates it from fail2ban for web-facing operators. Under the hood it embeds Coraza, a Go-native implementation of the OWASP Core Rule Set — the same battle-tested rulesets commercial WAFs use to block SQL injection, XSS, path traversal, and the rest of the OWASP Top 10. You can run CrowdSec as a WAF in front of your web apps, not just an SSH brute-force blocker.
This matters because most self-hosted stacks expose a web UI — a dashboard, a admin panel, a Nextcloud — and the attacks against those are application-layer, not just network-layer. A firewall ban stops a port scan; it does not stop a crafted request that exploits a vulnerable endpoint. The WAF layer inspects the request itself against the CRS and blocks malicious payloads. For an operator running several web services behind a reverse proxy, having one engine that does both network and application protection, configured in one place, is a meaningful simplification.
The WAF is more involved to tune than the basic IPS. The OWASP CRS is notorious for false positives on legit traffic if left at strict settings, so you typically start in detection-only mode, watch what it would block, and relax rules that bite real users. CrowdSec's blog has explicit guides for Traefik and Nginx virtual-patching setups, which tells you the project expects operators to do this work. It is powerful; it is not fire-and-forget. The sovereignty win is that you get enterprise-grade WAF logic without an enterprise WAF bill or a cloud dependency.
5. CrowdSec vs fail2ban: The Real Comparison
The honest comparison is fail2ban, because that is what CrowdSec replaces for most people. CrowdSec's own benchmark — 60x faster, IPv6-compatible, Go versus Python — is partly marketing but partly real: Go's concurrency model handles high log volumes better than fail2ban's Python loops, and IPv6 support is native where fail2ban's has historically lagged. For a busy edge node ingesting many megabytes of logs per minute, CrowdSec's performance edge is noticeable.
Functionally, the bigger differences are architectural. fail2ban is monolithic: it reads logs, decides, and writes iptables rules itself, per jail. CrowdSec separates detection (agent) from enforcement (bouncer), so one agent can protect many surfaces and the enforcement points are pluggable. fail2ban has no community intelligence — every instance relearns attacks alone. CrowdSec's network is its headline advantage and its only real caveat (the telemetry). If you want pure local-only and you are fine re-learning attacks, fail2ban is simpler. If you want community intelligence and a WAF, CrowdSec is the upgrade.
There is a migration story: many operators run both during a transition, or replace fail2ban entirely once CrowdSec's collections cover their services. The configuration is more involved upfront (enabling collections, installing bouncers) but pays off in coverage. For a homelab with a handful of exposed services, the setup is an evening; for a fleet, the Hub's collections and the single agent model scale better than per-host fail2ban jails.
6. Bouncers: Where Decisions Become Blocks
The bouncer is where CrowdSec's decisions turn into actual dropped packets or 403s, and choosing the right one for each surface is the operational core. The firewall bouncer (nftables/iptables) is the blunt instrument: banned IPs simply cannot reach your machine at the network layer. This is the right default for SSH and any service where a banned IP should be invisible. It is fast and cheap, but it is all-or-nothing per IP.
The web bouncers (Nginx, Traefik) operate at the proxy and can do more nuanced things: return a 403, or present a Captcha, or challenge a suspected bot while letting humans through. For a public web app where you do not want to hard-block a whole IP that might be a shared NAT, the Captcha challenge is the gentler enforcement. Cloudflare and other CDN bouncers push decisions to the edge provider, useful if your traffic rides through them.
Installing a bouncer means giving it access to the agent's LAPI (a local API key), so the bouncer can pull current decisions. The decoupling means you can add a bouncer later without touching detection, or run several bouncers off one agent. The operational lesson: think in terms of "what do I want to happen to a banned IP at each surface," then pick the bouncer that does it. Most stacks end up with a firewall bouncer plus one web bouncer, which covers the common cases.
7. What It Honestly Costs
CrowdSec the engine is MIT-licensed and free, with no per-seat or per-node fee. Running it costs the same as any lightweight Go daemon: a few tens of MB of RAM, negligible CPU at rest, and a small amount of disk for the local decision database. On a homelab node it is background noise. The standalone detection + bouncer setup is effectively free infrastructure.
The cost appears if you adopt the CrowdSec Console — the SaaS dashboard that aggregates signals across your instances, shows attack maps, and centralizes management. The Console is a commercial offering with tiers, and while the local engine never requires it, the polished multi-instance view does. For a single node, you do not need it;
cscli and the local decisions list are enough. For a fleet, the Console's convenience has a price, and that is the project's sustainable-business line. Name it: the software is free, the management UX at scale is not, and that is a fair open-core-adjacent model — except the core engine itself is fully MIT, not a teaser.
The other cost is tuning time. The WAF especially wants attention: watch detection-only mode, tune false positives, then enforce. The Hub collections want curation: enable what you run, ignore what you do not. This is the normal security-maintenance tax, and for someone who already runs a server it is manageable. For someone who wants "install and forget," CrowdSec will drift out of tune if you never look at it — like any security tool.
8. Integrating With Your Existing Stack
CrowdSec slots into a self-hosted stack cleanly because it speaks the languages your other tools already use. Behind a Traefik or Nginx reverse proxy (both covered in this series), the web bouncer protects everything the proxy fronts — which is usually your entire app surface — from one place. Pair it with Caddy or Traefik for TLS, and you have a protected, encrypted edge with one detection engine.
For identity, CrowdSec complements rather than replaces your SSO. Authelia or Authentik (covered earlier) decides who is authorized; CrowdSec decides which IPs are too hostile to even get a login prompt. They are different layers — one is authentication, the other is pre-auth threat blocking — and running both means an attacker has to survive CrowdSec's bans before they ever reach Authelia's login form. That layering is good defense-in-depth and a natural fit for the sovereign stack this blog builds.
There is also an MCP (Model Context Protocol) server in the CrowdSec ecosystem now, letting AI tooling write WAF rules and scenarios against your instance with natural language. That is early and you should treat it as experimental, but it signals where the project is heading: making detection-rule authoring accessible to people who do not want to hand-write expressions. For an operator who spends time in agentic tooling, it is worth watching.
9. The Privacy Story, Stated Plainly
Because this blog is about data sovereignty, the telemetry question deserves a direct answer. In default participation mode, CrowdSec sends: attacker source IPs and the attack type (e.g., "SSH brute force from 1.2.3.4"). It does not send your logs, your users, your content, your internal IPs, or anything about what you run beyond the signal that an attack happened. The stream is narrow and low-sensitivity by design — it is the hostile party's address, not yours.
What you receive in return is the community blocklist: IPs the network has flagged, with reputation scores. Your agent uses these to pre-ban known-bad sources. Neither direction reveals your infrastructure's contents; the outbound direction reveals only "this IP attacked me," which is not sensitive information about you. If even that is too much — air-gapped, stealth-required, or simply "no outbound from my security layer" — standalone mode disables participation entirely and you keep purely local detection. The choice is a flag, not a fork.
The sovereignty framing: CrowdSec is a tool you run, under MIT, with source you can read. The only non-local behavior is the optional, configurable intelligence sharing, which is the feature that makes it better than fail2ban. An operator who wants the benefit without the outbound signal simply flips participation off. Few security tools make that trade this explicit, and it is worth crediting.
10. Operational Reality: Alerts, False Positives, and Tuning
Day-to-day, CrowdSec is mostly quiet — which is the goal. When it acts,
cscli decisions list tells you who and why. The first week of a new deployment is the tuning week: you watch what gets banned, confirm the bans are correct, and adjust scenario thresholds if a legit pattern is tripping them. A common early false positive is a script or a user behind a shared IP hitting rate limits; you whitelist those or raise thresholds.
The WAF needs the most care. The OWASP CRS will, in strict mode, block legitimate requests that merely look like injection attempts — a search box query with a
< character, for example. Run detection-only first, review the would-block log, and disable or relax the specific rules biting real traffic. CrowdSec's documentation and blog walk through this for the major proxies. The operator who treats the WAF as "watch, then enforce" has a clean experience; the one who flips it to block on day one gets support tickets from themselves.
Remediation when something goes wrong is straightforward:
cscli decisions delete clears a bad ban, or you ban an IP manually with cscli decisions add. Because decisions are time-boxed (a ban expires), a mistaken block self-heals if you do nothing — another design choice that favors "safe by default" over "permanent by accident." For incident response, the local decisions list is your first stop, and it is fast.
11. When CrowdSec Is the Wrong Tool
CrowdSec is not a full SIEM, not an EDR, and not a replacement for a coherent security posture. If you need centralized log analytics across a large org, enterprise SOC workflows, or endpoint detection on workstations, CrowdSec is one layer, not the stack. It detects and blocks at the network and application edge for the services it understands; it does not inventory your assets or hunt threats across a heterogeneous fleet by itself.
It also assumes you run services worth protecting and expose them. If your infrastructure is fully internal with no public surface, the attack traffic that feeds detection is minimal, and the value drops — though internal lateral-movement detection via internal log sources is possible, it is not the common deployment. And if your threat model is "a nation-state targeting me specifically," a shared community blocklist is a speed bump, not a wall; you need defense in depth CrowdSec cannot provide alone.
Finally, if you will never tune it, a WAF you ignore is a liability (false blocks) and an IPS you ignore is a fading shield. CrowdSec rewards the operator who checks it; it does not replace attention. That is true of all security tooling, but it is worth saying plainly before you install something that can block your own traffic.
12. The Sovereignty Verdict
CrowdSec is what modern, community-aware intrusion prevention looks like when the engine is MIT-licensed and the operator stays in control of the data flow. The detection is real, the WAF is enterprise-grade logic without the enterprise bill, and the architecture — agent, hub, bouncer — scales from one node to a fleet without a rewrite. The community network is a genuine advantage over the fail2ban island, and the fact that you can disable it entirely is the sovereignty guarantee.
For a self-hosted stack, CrowdSec belongs at the edge: behind your reverse proxy, in front of your SSO, banning the bots before they reach a login form. The cost is an evening to set up, occasional tuning, and a conscious decision about participation telemetry. The return is protection that gets smarter because the herd gets attacked — exactly the kind of leverage a solo operator wants. Run it, point a bouncer at your proxy, and your public IP stops being an easy target.
13. Installation: A Minimal Docker Compose That Actually Works
The supported way to run CrowdSec is the official image plus at least one bouncer. A minimal realistic stack has two containers: the
crowdsec agent and a firewall bouncer (for example crowdsec/cs-firewall-bouncer) that reads decisions and installs nftables rules. You mount your service logs into the agent — typically by sharing the host's /var/log read-only, or by pointing the agent at journald — and you enable the collections for what you run.
The configuration reads like this in practice: the agent container gets the log volumes and an environment that points at the LAPI; the bouncer container gets the agent's API key (generated on first boot via
cscli, then injected) and the host's nftables socket so it can write rules. For Docker-shaped stacks, CrowdSec publishes a ready docker-compose in its docs; the work is enabling collections (cscli collections install crowdsecurity/nginx crowdsecurity/sshd) and confirming the bouncer registers. First boot auto-configures baseline detection, which is why the project claims "functional out-of-the-box" — and for common services, that claim holds.
The one gotcha newcomers hit: the bouncer must reach the agent's LAPI, which means the two containers share a network or the agent exposes LAPI on a reachable address. If the bouncer logs "cannot connect to LAPI," nothing gets enforced even though detection is running. Verify with
cscli bouncers list — you should see your bouncer registered. That single check catches most "why isn't it blocking?" confusion.
14. Writing Your Own Scenario
The Hub covers mainstream services, but the moment you self-host something bespoke, you write a scenario, and it is easier than it sounds. A scenario is YAML: it names a filter (which parsed events it watches), an expression (the condition), a scope (what gets banned — usually the source IP), and a duration. A real example: "if more than 20 HTTP 404s from one IP in 30 seconds, ban for 1 hour." The expression language is a sandboxed evaluator, so you write
evt.Meta.log_type == 'http_access' && evt.Parsed.status == '404' and a count() > 20 over a window.
The power is in composing signals. You are not limited to one log source — a scenario can correlate an SSH failure with a web login attempt from the same IP and ban on the combined behavior, something fail2ban's per-jail model resists. Once written, you load it with
cscli scenarios install (or drop the file in the config directory) and it is live. The community Hub is itself mostly user-contributed scenarios, so the pattern scales: you benefit from others' detections and can contribute yours. For an operator who runs unusual software, this is the feature that makes CrowdSec fit any stack, not just the popular ones.
15. Community Engine vs the Console: Drawing the Line
It bears repeating where the free ends and the paid begins, because "open source security tool" often hides an open-core cliff. CrowdSec's detection engine, the Hub content, the bouncers, and the WAF are all MIT and free — there is no "enterprise edition of the engine" with the good parts locked away. The commercial product is the CrowdSec Console: a hosted dashboard that aggregates signals across your instances, shows attack maps, centralizes scenario management, and adds team features. You never need it for one node.
What the Console adds is operational visibility at scale, not detection capability. A single host gets everything from the free engine plus
cscli. A fleet of fifty hosts benefits from one pane of glass showing all decisions — and that pane is the paid line. This is a sustainable model done relatively honestly: the software is genuinely open, the convenience of centralized SaaS management is the product. If your sovereignty rule is "no SaaS in the security path," you simply never enable the Console and lose nothing functional. That clean separation is rarer than it should be and worth crediting.
16. Integrating With a CDN: The Cloudflare Bouncer
If your traffic rides through Cloudflare or a similar CDN, the naive firewall bouncer is less effective, because the edge sees Cloudflare's IPs, not the attacker's. CrowdSec solves this with a Cloudflare bouncer that pushes decisions to Cloudflare's firewall API, so a banned real IP is blocked at the CDN before it reaches you. This matters because for many self-hosters, Cloudflare is the actual edge, and banning at your origin only after Cloudflare forwarded the request leaves a window.
The same pattern exists for other CDN and cloud WAF providers. The mechanism is the same as any bouncer: the agent emits a decision, the bouncer translates it into the provider's block API. The trade is that the bouncer now holds API credentials to your CDN, which is a trust relationship you should scope tightly (a token that can only manage firewall rules, not account-wide). For an operator already behind Cloudflare, the CDN bouncer is the correct enforcement point and closes the "I banned them but Cloudflare forwarded it anyway" gap that surprises people new to edge architecture.
17. A Real Attack Posture: What Gets Blocked in Week One
To make the value concrete, here is what a typical public home server sees in its first week with CrowdSec on. SSH: a steady drizzle of brute-force attempts from credential-stuffing bots, each banned after a handful of failed logins — exactly the fail2ban job, now shared with the network so the same bot is pre-banned next time. Web: scans for
/wp-admin, /.env, /phpmyadmin, and a hundred other fingerprints of software you do not run; the HTTP collections flag and block the probing IPs.
More interesting are the application-layer hits the WAF catches: attempted SQL injection in a search parameter, path traversal (
../../etc/passwd), and XSS payloads in form fields. These never reach your app because the CRS blocks them at the proxy. The decisions list becomes a mini threat feed for your own infrastructure — you can see, in real time, that someone is probing for a specific CVE, and decide whether your stack is exposed. For an operator who likes to know what the internet is throwing at their box, this visibility is half the value; the blocking is the other half.
18. Observability: Metrics and Grafana
CrowdSec exposes Prometheus metrics, which means it drops into a observability stack cleanly. If you already run Grafana (covered earlier in this series) and scrape your homelab, you can plot bans over time, top attacked services, and decision volume on the same dashboard as your CPU and uptime. This turns security from a silent background process into a visible signal you actually watch — which is the difference between a tool you maintain and a tool you forget.
The metric that matters most operationally is "decisions per service per day." A sudden spike tells you a new campaign found your IP; a drop to zero on a service that should see traffic might mean a parser stopped parsing (a log path changed, a collection got disabled). Wiring CrowdSec into Grafana turns those anomalies into dashboard lines you will notice, rather than a log file you will never read. For a sovereignty-minded operator who already monitors their stack, this integration is the natural completion: protect the edge, then watch the protection work.
19. Upgrading and Staying Current
CrowdSec ships releases steadily — the reviewed line is v1.7.8 (a security release fixing a WAF bypass and a LAPI DoS), with v1.8.0 in release candidate during 2026. Upgrades are routine container pulls, but the discipline that prevents surprises is the same as everywhere: read the release notes, because a security release may change a default or a bouncer protocol. The bouncer and agent versions should stay compatible; mismatches usually still work but occasionally surface enforcement gaps.
Because detection and remediation are decoupled, an agent upgrade does not require rebuilding your firewall rules — the bouncer keeps enforcing the last decisions until it reconnects. Take a config backup before a major bump (the
cscli config and your custom scenarios), and roll forward. The WAF ruleset (Coraza/CRS) updates on its own cadence; track those separately if you run AppSec, since a CRS update can change false-positive behavior and deserves a detection-only pass after applying. As with every tool in this series, the operator who plans the upgrade has a boring life; the one who blind-pulls at 3 a.m. occasionally does not.
20. Troubleshooting: The Three Questions That Fix Most Problems
"When it doesn't block, what's wrong?" Almost always one of three things. First, the bouncer is not registered — check
cscli bouncers list; if it is empty, the bouncer never connected to LAPI, so fix the network or the API key. Second, the relevant collection is not installed — if you run Traefik but only enabled the nginx collection, Traefik logs are unparsed and nothing detects them; install the right collection. Third, the log source is not reaching the agent — a wrong volume mount or a journald permission means the agent sees no events and has nothing to decide on.
"If it blocks too much," you have a false positive in a scenario or a WAF rule. Find the decision (
cscli decisions list), identify the scenario that emitted it, and either raise its threshold, whitelist the IP, or switch the WAF rule to detection-only while you tune. "If the Console shows different data than local," you have a sync issue with the SaaS — but for a standalone operator this question never arises, because there is no Console. The pattern is reassuring: CrowdSec's failures are observable and local, and every one of them is fixable from cscli without calling a vendor, because there is no vendor in the loop. Run it, watch it for a week, and the internet's noise becomes a list you control — and the herd makes that list smarter the longer you stay in it.
Related
For the rest of a sovereign, self-hosted stack, these pieces from our series travel with CrowdSec:
- Caddy: The Web Server That Encrypts for You — the reverse proxy to run the CrowdSec web bouncer behind, for TLS plus AppSec in one place.
- Authelia: SSO and 2FA for Everything Behind Your Reverse Proxy — the auth layer CrowdSec protects; banned IPs never reach its login form.
- Traefik: The Edge Router for Container Stacks — another proxy option with a mature CrowdSec middleware pattern.
- Keycloak: The Open-Source Identity Platform With No Open-Core Trap — the heavier IdP to pair with edge protection for full identity-plus-threat layering.
- Uptime Kuma: Know When Your Stack Is Down — monitor that your security layer and the services it protects are actually up.
Comments (0)
No comments yet. Be the first to comment!