1. What Exactly Is Netdata?
Most monitoring tools ask you to make a quiet trade. Give us a query language, give us a cardinality budget, give us a few hundred megabytes of RAM, and in return we will tell you what happened to your servers — roughly, in fifteen-second intervals, after the data has already taken a round trip to someone else's datacenter. Netdata, from the project of the same name, declines that trade in the most emphatic way possible: it collects every metric, every second, with zero configuration, and it does the analysis on the box that produced the data. There is no sampling, no mandatory cloud account, and no default behavior that ships your telemetry anywhere you did not explicitly choose.
At the time of writing the repository sits at roughly 80,000 GitHub stars, making it one of the most popular open-source observability projects in existence. It is written primarily in C for the core daemon, with Go and Rust handling the plugin and collector layers. The license is GPL-3.0 for the agent, which means the core is genuinely free and the copyleft is real — anyone distributing a modified Netdata must also release their changes under the same terms. That matters more than it sounds for a monitoring tool, because monitoring agents run with elevated privileges and see everything your systems are doing.
Netdata positions itself as "edge-native observability." The phrase is marketing, but it describes something concrete: the agent is designed to live on the machine it observes, to stay lightweight, and to make decisions locally. The company behind it publishes twelve "principles," one of which is explicitly "data sovereignty by design." This is not a bolt-on feature. The architecture assumes your metrics stay on your infrastructure unless you deliberately bridge them to Netdata Cloud, and even then the raw time-series data does not leave the node — only control messages and metadata cross the link.
For a self-hoster, that framing is the entire point. You are not renting the ability to see your own machines. You own the collector, the storage, the dashboards, and the alerts. The only thing Netdata the company can sell you is convenience, not your data.
2. Why Data Sovereignty Matters for Monitoring
It is easy to underestimate how much a metrics stream reveals. A per-second CPU chart tells an attacker when your backup jobs run. A memory graph leaks the size of your in-memory caches. Network flow data exposes which internal services talk to which, and at what volume. Disk I/O patterns can betray the database engine you thought was hidden behind an API. In a very real sense, your observability pipeline is a high-fidelity side channel into your entire operation.
This is why the data-sovereignty question is not academic for monitoring the way it sometimes is for, say, a static website. When you point a SaaS monitoring agent at your infrastructure, you are handing a third party a continuous, structured, timestamped record of your system's behavior — often including hostnames, internal IP addresses, and service topology. Depending on the vendor's jurisdiction, that record may be reachable under foreign surveillance law, retained longer than you would like, or aggregated into models you never agreed to.
Netdata's answer is architectural rather than contractual. Because the agent stores and analyzes data locally, the blast radius of a Netdata compromise is limited to the node it runs on, not your entire fleet's history in someone's data lake. You can run Netdata fully air-gapped. You can refuse the cloud link entirely and lose nothing from the core experience. And because the storage engine is local and queryable, you retain direct access to the raw numbers in a way that SaaS dashboards frequently obscure behind proprietary APIs.
This is the same sovereignty argument that pulls people toward self-hosted analytics or a private photo library, except the stakes are arguably higher because monitoring sees everything else. If you have already decided that your documents, your media, and your identity provider belong on hardware you control, it is inconsistent to then ship the telemetry that describes all of it to a third party by default.
3. Under the Hood: A C Core Wrapped in Go and Rust
The first thing worth understanding about Netdata is that it is not a single-language application wearing a uniform. It is a C core surrounded by satellites written in Go and Rust, stitched together through a well-defined plugin protocol. This is a deliberate performance decision, not legacy baggage.
The C core — the daemon itself, the time-series database, the health (alerting) engine, the exporting layer, and the agent-cloud link — handles the hot path: ingesting millions of samples per second with minimal allocations and direct control over memory layout and I/O. The Go layer,
go.d.plugin, is the workhorse collector for application-level integrations: databases, message brokers, web servers, DNS, and hundreds of other targets, including Prometheus-compatible scraping. Go is the right tool there because goroutine-per-job concurrency and fast HTTP client work make it easy to add integrations without recompiling the C core. The Rust components, living under
src/crates/, are the emerging layer for memory-safe data processing and ML workloads.
One detail that separates Netdata from many self-hosted options is its eBPF plugin. Rather than reading only from
/proc, Netdata can attach eBPF programs directly to kernel probes to collect metrics that are simply not available any other way — TCP retransmits per process, file descriptor lifetimes, VFS call counts. This is genuinely advanced instrumentation, and it is shipped in the standard Docker image.
Privilege separation is handled through a plugin model rather than running the whole daemon as root. Individual plugin binaries carry the setuid bit and are launched by the unprivileged
netdata user over a pipe-based protocol. This is the correct way to do elevated collection, and it is reassuring to see it enforced by default rather than bolted on after a CVE.
4. Zero-Configuration Auto-Discovery in Practice
The single feature that makes Netdata spread through homelabs and small ops teams is that it works before you configure anything. Install the agent, open the dashboard, and within seconds you are looking at CPU, memory, disk, network, processes, containers, and dozens of application metrics that were detected automatically. There is no scrape config to write, no exporter to deploy, no service discovery YAML to argue with.
In practice this means a fresh Ubuntu or Debian box with a one-line kickstart script will, within a minute, show you per-application breakdowns, systemd service health, and even the energy estimate for what the machine is drawing. On a Docker host it discovers containers without you pointing it at the Docker socket manually in most setups. On a Kubernetes node it can map which pods and workloads talk to each other by reading the kernel rather than requiring you to instrument your applications.
A concrete example makes the value obvious. Take a small Docker host running a reverse proxy, a database, a cache, and two web apps. Before Netdata, answering "which container is eating my I/O at 3 a.m." means ssh-ing in, running
docker stats, and hoping the spike is still happening. After Netdata, the same question is a click: the per-container charts are already there, labeled, and second-resolution, and you can scrub back to exactly 3 a.m. yesterday. The discovery did not require you to tell Netdata that those containers exist; it read them from the runtime. The same is true for systemd services, network interfaces, and mounted filesystems. The agent is effectively a curious observer that maps your system faster than you could document it, which is precisely why it spreads through homelabs — the first five minutes are the demo.
The honest caveat is that "zero configuration" describes the collection, not the meaning. Netdata will happily show you a chart titled
apps.cpu with forty processes on it, but turning that into "is my PostgreSQL query workload healthy" still requires you to understand your own system. Auto-discovery removes the plumbing; it does not remove the judgment. For a solo operator or a small team without a dedicated SRE, that trade is almost always worth it, because the alternative — standing up a scrape-based pipeline and wiring exporters — is a project in itself before you have seen a single useful chart.
5. Per-Second Metrics and the DBENGINE Storage Engine
Netdata's headline claim is "every metric, every second, no sampling." That is not hyperbole, and it is made possible by a custom storage layer called DBENGINE rather than by throwing RAM at the problem. Traditional monitoring systems downsample or aggregate because storing raw high-frequency data is expensive. Netdata instead compresses samples aggressively — on the order of a fraction of a byte per sample for many metrics — and tiers the data between RAM and disk so that recent, hot data is fast while older data is cheap.
The practical consequence is that you can actually answer "what happened at 14:32:07" rather than "what was the fifteen-second average around 14:32." For troubleshooting a transient spike, a brief lock contention, or a one-second network blip, that resolution is the difference between finding the cause and staring at a smooth line that hid it. Per-second resolution also makes the on-device anomaly detection meaningful: a model trained on second-by-second behavior can flag a deviation that a downsampled series would average away.
Storage, of course, is not free. Per-second data accumulates. Netdata handles this with configurable retention and tiered storage, and the recommended approach is to set a retention window that matches your needs — often a few weeks of high-resolution data plus longer-term lower-resolution archives — rather than attempting to keep everything forever on a Raspberry Pi. For most self-hosters, a sensible retention policy keeps the disk footprint modest while preserving the diagnostic value that drew them to Netdata in the first place.
6. On-Device Machine Learning and Anomaly Detection
Perhaps the most misunderstood part of Netdata is its machine learning. This is not a cloud model that ingests your data to "help" you. The ML runs unsupervised, on the device, training per-metric models on the local baseline and then scoring new samples for anomaly. The claim is that you get "ML on every metric" without sending anything off-box and without standing up a separate ML platform.
In real use this is genuinely useful for surfaces you are not watching. A CPU chart you never open can still raise a flag when its behavior diverges from its own history. The models are trained locally and require a baseline period of collection before they become reliable, so a brand-new node will be quiet for a while — that is expected, not broken. The detection is per-metric and unsupervised, which means it is good at "this is weird relative to itself" and not at "this matches a known attack signature." Treat it as a tireless first responder that pages you to look, not as a security analyst that tells you why.
The data-sovereignty angle here is clean: because the training and inference happen at the edge, your metric values never need to leave the node to benefit from the model. That is a meaningful contrast with vendors who offer "AI insights" only when your data is in their cloud.
7. Alerting and Notifications Without a Sidecar
Netdata's health engine lives inside the agent. You define alert conditions — threshold breaches, rate-of-change, anomaly scores, absence of expected data — and the agent evaluates them locally and routes notifications through a long list of integrations: email, Slack, Discord, Telegram, Pushover, webhooks, and many more. There is no separate alerting service to deploy and no dependency on an external evaluation loop.
For a self-hoster this collapses a chunk of operational complexity. You do not need to run an additional rules engine or a notification broker just to be told that a disk is filling up or a service stopped responding. The alerting is part of the same daemon that collects the data, which keeps the configuration in one place and the failure modes simple.
The limitation is that alerting is per-node by default. If you want fleet-wide silencing, maintenance windows, and routing logic across hundreds of nodes, you will either connect to Netdata Cloud or build your own aggregation. For a single home server or a small cluster, the built-in engine is more than enough; for a large multi-tenant fleet, the local-first model starts to show its seams and you will want the parent-child centralization or an external layer.
8. The v2.11.0 Leap: Network, Logs, and OpenTelemetry
The release that put Netdata back in the headlines is v2.11.0, shipped in August 2026. It is a genuinely large jump rather than an incremental point release, and it pushes Netdata past the edge of the box it monitors and into the network connecting your boxes.
The headline additions are a Network Monitor dashboard, NetFlow/IPFIX/sFlow ingestion, SNMP-driven topology mapping, and an SNMP trap listener. In plain terms, Netdata can now watch your routers, switches, and firewalls the same way it already watched your Linux hosts — still self-hosted, still zero-config by default. For homelabbers and small ops teams who historically bolted on a separate tool just to see what was happening on the wire, this release aims to make Netdata the one dashboard that covers servers and network gear alike.
Alongside that, logs graduate into a first-class citizen. The Logs tab, previously limited to systemd journal and Windows events, now accepts OpenTelemetry log records directly, plus journald, filelog, and syslog receivers, configurable retention, and log-to-metric conversion. The same fast faceted query engine that serves metrics now also serves logs. An AWS CloudWatch collector rounds out the hybrid-cloud story, letting self-hosters fold cloud infrastructure into the same dashboards as their on-prem fleet.
These are optional capabilities. None of them require the cloud, and none of them turn on data egress by default. The Netdata Cloud additions in this release — AI-assisted troubleshooting and MCP support for connected nodes — are explicitly opt-in and do not affect the fully self-hosted workflow.
9. Installing Netdata on Your Own Hardware
Self-hosting Netdata is about as easy as it gets in this category. The official kickstart script detects your distribution, installs the agent, and enables it as a systemd service. A Docker deployment is a single container plus a handful of mounted volumes for configuration and the time-series database. Kubernetes users get a helm chart and the parent-child model for aggregating many agents under a single parent node.
A minimal Docker Compose for a single host looks roughly like this: you map the Netdata port, mount
/proc,
/sys, and
/var/run/docker.sock read-only where needed, set the capability for the eBPF and process plugins, and persist the
netdataconfig and
netdatalib volumes so your configuration and history survive container recreation. On a Raspberry Pi the same multi-arch image runs unchanged. The project documents macOS, Windows (via WSL or Docker), FreeBSD, and Linux endpoints, so a mixed fleet is covered.
The one operational habit worth forming early is scheduling archiving and, if you run a parent node, the centralization. Netdata separates raw collection from aggregation, and on larger setups you shift the expensive aggregation off the request path with a cron or systemd timer. Underprovisioned archiving is the most common cause of "my reports look stale," and it is trivially fixed once you know to look.
10. The Real Cost of Self-Hosting Netdata
"Free software" is never free to run, and monitoring is no exception. Here are the real numbers a self-hoster should expect, based on documented footprints and community reporting rather than vendor optimism.
Memory. A single Netdata agent on a small Linux host typically uses somewhere between 100 MB and 400 MB of RAM depending on how many collectors are active and how many applications it is tracking. The source-directory-style estimates put a minimum around 640 MB and recommend 2 GB headroom for comfortable operation, but those figures include the web server and a generous margin. On a dedicated monitoring parent node aggregating many children, budget 1–2 GB and scale with the number of nodes.
CPU. The agent is designed to be cheap — usually well under 5% of a single core on a lightly loaded host, spiking during archiving. The eBPF and process plugins add a little, but the design goal is "you should not be able to tell Netdata is running." On a Raspberry Pi 4 this holds; on a Pi Zero it gets tighter, and you would limit collectors.
Disk. This is the line item people forget. Per-second metrics with aggressive compression still accumulate. A single busy host can write several gigabytes a month at full resolution. With a tiered retention policy — say two weeks of high-resolution data plus longer low-resolution archives — a typical home server lands in the 10–30 GB range. Unsupervised, "keep everything forever" can quietly eat a 128 GB SD card.
Energy. The indirect cost is power. A always-on mini-PC or old laptop running your monitoring 24/7 might draw 10–25 W continuously. At an average electricity price, that is on the order of $1–4 per month for the hardware that hosts Netdata itself, before you count the machines it monitors. On a Pi, call it under a dollar a year in electricity.
A worked monthly example. Consider one always-on mini-PC (an older Intel NUC-class box, ~15 W draw) that hosts Netdata plus a handful of containers. Electricity at $0.15 per kWh: 15 W is 0.36 kWh per day, about 10.8 kWh per month, roughly $1.62 in power. The hardware was already owned, so capital is zero on the margin. Disk for two weeks of high-resolution plus longer low-resolution archives lands near 15 GB, which is nothing on a 500 GB SSD. If instead you had no always-on box and bought a Raspberry Pi 4 with a 32 GB card and a tiny power supply, the one-time cost is on the order of $60–80 and the continuous draw is under 4 W, so the electricity is pennies a year. The software license line is $0. The only recurring cost is your attention when an alert fires.
What it replaces, and what it does not. Against a SaaS monitor you pay nothing in license for the agent but you pay in hardware you already own or must buy, plus your own time. Against standing up Prometheus plus Grafana plus exporters plus an alertmanager, Netdata usually wins on setup time and loses on long-term flexibility. The honest cost summary: if you already own always-on hardware, Netdata is effectively free to run; if you would buy a box just to monitor other boxes, that capital and power cost is real and should be counted. The trap to avoid is treating "free agent" as "free total cost of ownership" — the disk and the always-on power are real, just small, and they are yours to size rather than a vendor's to bill.
11. Netdata vs. Grafana, Prometheus, and VictoriaMetrics
No self-hosting decision happens in a vacuum, and monitoring is the most crowded shelf in the self-hosted pantry. The relevant comparisons are Grafana, Prometheus, and VictoriaMetrics — all of which appear elsewhere in this blog's catalog.
Prometheus is the scrape-and-store standard. It is enormously flexible and integrates with everything, but it is a project to operate: you write scrape configs, manage retention, and usually pair it with Grafana for visualization and Alertmanager for notifications. Netdata is the opposite philosophy — collect automatically, store locally, visualize immediately. They are not mutually exclusive; many people run Netdata on each node for instant per-second visibility and Prometheus at the cluster level for long-term trends.
Grafana is a visualization layer, not a collector. Comparing it to Netdata is partly a category error, but in practice people reach for Grafana when they want dashboards, so it is the thing Netdata's own UI is measured against. Netdata's interface is purpose-built and fast for real-time drilling; Grafana is unbeatable for bespoke dashboards assembled from many sources. The two cooperate well: Netdata can export to Prometheus, and Grafana can display Netdata data.
VictoriaMetrics is a long-term time-series database that is Prometheus-compatible and exceptionally efficient at scale. It solves the retention-and-cost problem that raw per-second storage creates. Netdata solves the collection-and-instant-visibility problem. A common mature architecture is Netdata on every node, Prometheus or VictoriaMetrics as the long-term store, and Grafana as the presentation — though that is more moving parts than a solo operator needs on day one.
The short version: choose Netdata when you want to see everything now with minimal effort and keep the data local; choose the Prometheus/Grafana/VictoriaMetrics stack when you want maximum control, custom dashboards, and planet-scale retention, and you are willing to assemble and maintain it.
12. Where Your Data Goes: Data Sovereignty by Design
This is the section the whole review bends toward, so it deserves precision rather than slogan-repeat.
By default, nowhere but your disk. A standalone Netdata agent collects, stores, and queries data entirely on the host. No account is required. No telemetry about your telemetry is sent. The raw time-series never leaves the node.
The optional cloud link. Netdata Cloud is an opt-in control plane. When you connect a node, the Agent-Cloud Link carries control messages and metadata — things like node identity and dashboard state — but the architecture is explicitly designed so the raw metrics do not stream to the cloud. You can run for years with the cloud feature disabled and lose no core functionality. If your threat model forbids any outbound connection, you can firewall the agent and it will keep working locally.
Jurisdiction. The company is US-headquartered, which is worth naming plainly because sovereignty is partly about legal exposure. The mitigation is architectural: because your data is on your infrastructure, a subpoena or foreign-surveillance reach to Netdata the company does not automatically grant access to your metrics the way it would for a SaaS that stores them. The risk shifts to your own infrastructure security, which is the trade you accept with any self-hosted tool.
Retention and deletion. Because storage is local and queryable, you control how long data lives and when it is wiped. There is no vendor retention policy sitting between you and a right-to-erasure request. For regulated environments this direct control is a feature SaaS dashboards rarely offer without enterprise contracts. If a customer or colleague asks you to delete their associated telemetry, you can locate and remove it on your own disk on your own schedule, rather than filing a support ticket and hoping the vendor's retention job runs. That autonomy is the quiet dividend of self-hosting: compliance work becomes a filesystem operation instead of a vendor negotiation.
The one honest asterisk. Some advanced features — fleet management, RBAC, and horizontal scaling across very large deployments — are delivered through Netdata Cloud under commercial terms. The agent itself remains GPL-3.0 and fully functional. So the sovereignty story is clean for the core and conditional for the convenience layer, which is the right place to draw the line.
13. Honest Limitations You Should Know Before You Commit
No tool this popular is without seams, and the trustworthy review names them.
Long-term storage is not its strength. Netdata is built for high-resolution, recent data and instant query. If your primary need is years of retained history with cheap queries, a dedicated long-term store like VictoriaMetrics will serve you better, and you should plan to export rather than rely on Netdata alone.
Cardinality at scale. Per-second collection on a host with thousands of ephemeral containers or metrics can create high cardinality that stresses the local database. The project has improved this, but a single agent is not a replacement for a purpose-built cluster-scale TSDB.
Dashboard customization lags Grafana. The Netdata UI is excellent for exploration and real-time drill-down and weak for the elaborate, shared, branded dashboards Grafana users build. If bespoke dashboards are the point, Netdata is the wrong primary tool.
Learning the health language. Writing your own alert expressions takes familiarity with Netdata's configuration format. It is approachable, but it is not as universally documented as PromQL, simply because fewer people use it.
License nuance. The agent is GPL-3.0, which is friendly for personal and internal use but imposes source-disclosure obligations if you distribute a modified version. The UI carries an additional license (NCUL) that the project acknowledges needs clarification in places. For personal self-hosting this is a non-issue; for shipping a product that embeds Netdata, get legal eyes on it.
It sees a lot. Because the agent runs with elevated privileges and eBPF access, a compromised Netdata node is a privileged vantage point. The privilege-separation model mitigates this, but you should still treat the agent as a security-sensitive component: keep it patched, restrict who can reach its port, and put it behind your reverse proxy and auth like anything else on your network.
14. Hardening Netdata: Treating the Agent as a Privileged Component
Because Netdata collects at the kernel level and runs plugins with elevated permissions, the responsible thing is to treat it as a security-sensitive service rather than a harmless dashboard. The good news is that the project's own architecture makes hardening straightforward; the bad news is that the defaults prioritize convenience, so a little intent goes a long way.
The first move is to never expose the Netdata port directly to the internet. The dashboard is powerful and, by default, not protected by a login on a fresh install beyond what your network provides. The correct pattern is to place Netdata behind your reverse proxy — Traefik or Caddy both appear elsewhere in this catalog — terminate TLS there, and require authentication. Netdata supports its own simple password protection via the
web config, but pairing it with your existing single sign-on through Authelia or Authentik is the cleaner long-term answer and keeps one auth story for your whole stack.
The second move is to constrain the eBPF and setuid plugins to what you actually need. If you are monitoring a container host and have no use for kernel-level TCP retransmit metrics, you can disable the eBPF collector and reduce the privileged surface without losing the charts most people open. Every plugin you turn off is a piece of elevated code that cannot be abused. Netdata's modularity means this is a config change, not a rebuild.
The third move is network segmentation. A monitoring agent that can see every container and process is a high-value target, so it belongs on a management VLAN or at least behind the same firewall rules you would apply to your database. If you run a parent node that aggregates children, lock the parent-child communication to your internal ranges and monitor the monitor — an unavailable parent silently stops collecting, and you only notice when an alert you expected never arrives.
The fourth move is patching discipline. Netdata ships frequently — the v2.11.0 line alone shows a steady cadence — and the security hardening that lands in those releases matters for a privileged daemon. Pin your image to a specific digest rather than
latest so an update cannot silently change behavior, and read the release notes when you bump versions. The breaking change in v2.11.0 that ended support for RHEL 7, CentOS 7, and Amazon Linux 2 is a reminder that "self-hosted" also means "you own the upgrade decision."
Finally, consider the trust boundary of the cloud link even though it is optional. If you enable Netdata Cloud, understand precisely what crosses the wire: control messages and metadata, not your raw metrics, according to the architecture. If your threat model treats any outbound connection from a privileged host as unacceptable, disable the Agent-Cloud Link entirely and verify with outbound firewall rules. The point of sovereignty is that you get to decide, and Netdata gives you the switch — but only if you flip it deliberately rather than accepting the default.
15. Related
If Netdata scratches your itch for infrastructure you fully control, these other self-hosted pieces in this blog cover the surrounding territory:
16. Bottom Line
Netdata is the closest thing the self-hosting world has to "install it and immediately understand your machines," and it earns its 80,000 stars by refusing the usual monitoring compromise. Every metric, every second, analyzed on your hardware, with your raw data staying on your disk unless you explicitly bridge it elsewhere. For a solo operator, a homelab, or a small team that wants real-time visibility without standing up a scrape pipeline and a dashboard factory, it is hard to beat on effort-to-insight.
The trade you accept is that Netdata is a brilliant collector and real-time explorer, not a long-term warehouse or a dashboard studio. Plan to export to a dedicated time-series database if you need years of history, and do not expect it to replace Grafana for elaborate shared dashboards. The licensing is clean for personal use and worth a legal glance only if you intend to redistribute a modified agent. Treat the agent as the privileged component it is, keep it behind your auth and reverse proxy, and it will reward you with a level of per-second clarity that most paid monitors still only approximate.
If your rule is that your documents, your media, and your identity belong on hardware you control, your monitoring telemetry deserves the same standard — and Netdata is the most direct, most popular, and most sovereign way to enforce it. Spin it up on a spare machine this afternoon and you will understand your infrastructure better by dinnertime than most teams do after a quarter of SaaS dashboards.
Comments (0)
No comments yet. Be the first to comment!