Grafana Cloud

grafana.comcontributed by Samuele Ongaro

YES

Replaceable in one session with an AI coding agent.

Hosted Grafana with metrics, logs and traces attached: the dashboards you would run yourself, plus the storage behind them, operated by someone else.

Promptfree, for everyone, and the only version there is
Build me the observability stack I actually need instead of paying for Grafana Cloud: metrics, logs and dashboards, on my own machine, at a size I can operate.

Read this first: every piece here is free and open source. What the fee buys is somebody else running Prometheus, Loki and the storage behind them, and keeping the alerting up when your own infrastructure is the thing having a bad day. That last point is the real one — read the section on it before deciding.

STACK
- Prometheus for metrics, Loki for logs, Grafana for dashboards. Do not rewrite any of them
- Node 20+ with Fastify only for the small pieces you need around them
- Docker or Podman, on a machine that is not the one being watched
- Caddy in front

WHAT TO ACTUALLY RUN
- Prometheus with a retention of a few weeks on local disk, which is plenty for the questions you will ask
- Loki with the filesystem backend and a stated retention, or an S3-compatible target if the volume grows. Loki indexes labels only, not log content, which is why it is cheap — keep the label cardinality low and it will run happily on very little
- Grafana with its own SQLite database, provisioned from files
- Alertmanager for routing, deduplication, grouping and silences. Do not write this yourself: grouping and inhibition are subtle and it is already solved
- node_exporter on every machine, and whatever exporter matches each service

INSTRUMENTING YOUR OWN APPLICATION
- Four metric kinds and no more: a counter for events, a gauge for a current value, a histogram for durations, and a summary only if you know why you want one
- The four that answer almost everything: request rate, error rate, duration histogram, and saturation of whatever the bottleneck is
- Labels are dimensions, not identifiers. Never put a user id, a request id, an email or a full path with parameters in a label — each distinct value is a new time series, and a handful of unbounded labels will exhaust the memory of the machine. This is the single most common way a self-hosted metrics stack dies
- Histogram buckets chosen for the thing being measured, not left at the default, or every percentile will be a lie

LOGS
- Structured, one JSON object per line, with a level, a timestamp, a message and fields
- A request id in every line of a request, which is what makes a log searchable after the fact
- Labels on the stream — application, environment, host, level — and everything else in the line itself
- Secrets, tokens and personal data redacted at the point of logging, by key name and by pattern. A log is a place data leaks to and stays

DASHBOARDS
- Provisioned as JSON files in a repository, not clicked into existence and lost. A dashboard nobody can recreate is a dashboard that dies with the disk
- Three that matter: one per service with the four signals, one for the machine, and one overview showing whether anything is currently wrong
- Resist building forty. A wall of graphs nobody reads is worse than three that get looked at every morning

ALERTING, AND THE HONEST PROBLEM
- Alert on symptoms, not causes: users are getting errors, the queue is growing, the disk will be full in four hours. High CPU is not an alert
- Every alert has a written runbook link and a severity, and an alert with no action attached should be deleted
- For-durations on everything, so a thirty-second blip does not page anybody
- Here is the problem this stack cannot solve on its own: it runs on your infrastructure, so when your infrastructure is down it is down too, and the alert never arrives. Run it on a different machine at a different provider, and add one external check from a third party that tells you the whole thing is unreachable. That external check is the one piece you should not self-host

BACKUP AND RETENTION
- Metrics and logs are expendable; dashboards, alert rules and provisioning files are not, and those live in a git repository
- Grafana's own database backed up nightly
- Retention stated for each, and a disk watchdog, because filling the observability machine's disk is how you lose visibility during the incident that filled it

WHAT IT COSTS TO OPERATE
- Write down honestly: upgrades, disk management, cardinality accidents, and a stack that needs attention on the day everything else needs attention
- At a small scale this is a few euros a month and an afternoon a quarter. At a large scale it is a job, and the hosted product starts to look reasonable
- Say which side of that line you are on before starting

WHAT MATTERS MOST
Label cardinality and putting the stack somewhere else. Get instrumentation right at the start — a wrong label choice is discovered when the machine runs out of memory — and never host the thing that watches your servers on one of the servers it watches.

Give me the compose files, the Prometheus and Loki configuration, the Alertmanager routing, three provisioned dashboards, an example instrumented service, and a README that reads as a runbook including the cardinality rules and the external check.

What you lose

  • Running Prometheus, Loki and Tempo yourself, which is real operational work at any scale
  • Long retention without you managing the disks
  • Alerting that keeps working when your own infrastructure is the thing that broke

If you would rather not build

  • Netdata, for a single machine with no configuration
  • VictoriaMetrics, for cheaper long retention

What it costs

read from their page 15 Aug 2026

PlanBilled monthlyBilled yearlyLast read
—$19/mo—15 Aug 2026

Their pricing page is where these came from. Seeing a different price? Tell us.

The escape hatch

open source · no votes, no paid placement

Prometheus

$0

Metrics collection and alerting rules, the standard for this job.

prometheus/prometheusfree · open source

Grafana

$0

The dashboards themselves, free and self-hostable.

grafana/grafanafree · open source

Why this verdict

our own opinion · changed only by a person

70/100

Verdict yes at 70. Every component is free; you are paying for operation and retention. The alerting-elsewhere rule is the one thing not to compromise on.

History

tracked since 10 Aug 2026 · nothing is ever overwritten

Interest · last 30 dayspeak 2/day
views01230 Aug4 Sept9 Sept14 Sept19 Sept24 Sept28 Sept
— views— prompt copies none yet— votes none yet

Questions about Grafana Cloud

answered from the record above

Is Grafana Cloud free?

No — the plan we track is $19 a month. Pro from around $19/month billed monthly plus usage; a free tier covers small projects.

Can you replace Grafana Cloud by building your own?

YES. Replaceable in one session with an AI coding agent. Replacement score 70 out of 100, build time one session. Read what you lose before you decide.

How much does Grafana Cloud cost?

$19 a month on Pro — $228 a year. Recorded 10 Aug 2026.

What do you lose by replacing Grafana Cloud?

Running Prometheus, Loki and Tempo yourself, which is real operational work at any scale; Long retention without you managing the disks; Alerting that keeps working when your own infrastructure is the thing that broke. If any of those carry weight for you, keep paying.

Is there an open-source alternative to Grafana Cloud?

Yes: Prometheus, Grafana. The prompt on this page is for when you want it your way instead.

Related entries

same category first, most replaced first

All 25 in Monitoring & uptime

Not sending yet

Every week, something stops being worth paying for.

New verdicts, prices that moved, entries added. One email a week. Unsubscribe in one click. Nothing is being sent yet — your address is kept here, and the first issue is the first thing it is used for.

free forever · no tracking pixel · stored here, never passed to anyone

Esc