Master Grafana: Key Questions for Admins and Engineers

If you’re interviewing for a Grafana Administrator, Monitoring Engineer, SRE, DevOps Engineer, Platform Engineer, OpenShift Monitoring, or Cloud Architect role, these are some of the most common Grafana interview questions and answers.

1. What is Grafana?

Grafana is an open-source observability and visualization platform used to query, visualize, alert on, and analyze metrics, logs, and traces from multiple data sources.

Supported sources include:

  • Prometheus
  • Loki
  • Elasticsearch
  • InfluxDB
  • Graphite
  • OpenSearch
  • Azure Monitor
  • CloudWatch

2. Explain Grafana Architecture

Components

Data Sources

  • Prometheus
  • Loki
  • Elasticsearch
  • CloudWatch

Grafana Server

  • User authentication
  • Dashboard rendering
  • Alerting
  • API access

Database

  • SQLite (default)
  • MySQL
  • PostgreSQL

Users

  • View dashboards
  • Create dashboards
  • Manage alerts

3. What is the difference between Grafana and Prometheus?

GrafanaPrometheus
Visualization ToolMonitoring System
Creates DashboardsCollects Metrics
Alert VisualizationAlert Rules
Multiple Data SourcesPrimarily Metrics

Example:

Prometheus stores:

node_cpu_seconds_total

Grafana displays it as:

  • Graph
  • Gauge
  • Heatmap
  • Table

4. What are Grafana Data Sources?

A data source is the backend from which Grafana retrieves data.

Examples:

  • Prometheus
  • Loki
  • Elasticsearch
  • CloudWatch
  • Azure Monitor
  • SQL Databases

5. Explain Grafana Dashboards

A dashboard is a collection of panels displaying metrics.

Examples:

  • CPU Utilization
  • Memory Usage
  • Disk Usage
  • Network Traffic
  • Application Response Time

6. What are Panels in Grafana?

Panels are visualization widgets.

Common types:

  • Time Series
  • Gauge
  • Stat
  • Table
  • Pie Chart
  • Heatmap
  • Logs
  • Geomap

7. What are Variables?

Variables make dashboards dynamic.

Example:

$cluster
$namespace
$pod

Instead of creating 100 dashboards, create one dashboard and switch values.


8. What is PromQL?

PromQL is the query language used by Prometheus.

Example:

CPU Usage:

100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)

Memory Usage:

(node_memory_MemTotal_bytes-node_memory_MemAvailable_bytes)
/
node_memory_MemTotal_bytes *100

9. Explain Grafana Alerting

Grafana Alerting evaluates queries and generates notifications.

Flow:

Metric
Query
Condition
Alert
Notification

Example:

CPU > 80%
for 5 minutes

10. What Notification Channels are supported?

  • Email
  • Slack
  • Teams
  • PagerDuty
  • Webhook
  • Opsgenie

11. Explain Alert Rule Components

Query
node_load1
Condition
IS ABOVE 80
Evaluation
Every 1 minute
Alert State
  • OK
  • Pending
  • Alerting
  • No Data

12. Difference Between Loki and Elasticsearch

LokiElasticsearch
Log AggregationSearch Engine
Stores LabelsFull Indexing
Lower CostHigher Storage
Optimized for GrafanaGeneral Purpose

13. What is Loki?

Loki is Grafana’s log aggregation platform.

Components:

Promtail
Loki
Grafana

14. What is Promtail?

Promtail collects logs and forwards them to Loki.

Example:

scrape_configs:
- job_name: system
static_configs:
- targets:
- localhost
labels:
job: syslog
__path__: /var/log/*.log

15. Explain Grafana Mimir

Grafana Mimir is a scalable Prometheus-compatible metrics backend.

Benefits:

  • Multi-tenant
  • Long-term retention
  • HA
  • Massive scale

16. What is Grafana Tempo?

Grafana Tempo stores distributed traces.

Supports:

  • OpenTelemetry
  • Jaeger
  • Zipkin

17. Explain the LGTM Stack

Very common interview question.

L = Loki
G = Grafana
T = Tempo
M = Mimir

Together they provide:

  • Metrics
  • Logs
  • Traces
  • Dashboards
  • Alerting

18. What is OpenTelemetry?

OpenTelemetry is a framework for collecting:

  • Metrics
  • Logs
  • Traces

and exporting them to Grafana.


19. How do you secure Grafana?

Best practices:

  • SSO (SAML/OIDC)
  • RBAC
  • HTTPS
  • MFA
  • LDAP integration
  • Dashboard permissions
  • Network restrictions

20. Explain Grafana RBAC

Roles:

Viewer

Read only

Editor

Create dashboards

Admin

Manage users and settings


21. Grafana High Availability Design

Load Balancer
|
+-----+-----+
| |
Grafana1 Grafana2
| |
+-----+-----+
|
PostgreSQL

Interview answer:

Grafana servers are stateless. Store configuration in PostgreSQL and place multiple Grafana instances behind a load balancer.


22. Grafana Troubleshooting Questions

Dashboard Empty

Check:

  • Data source connectivity
  • Query syntax
  • Time range
  • Permissions
Alert Not Firing

Check:

  • Alert evaluation interval
  • Query results
  • Notification policy
  • Contact point
Slow Dashboard

Check:

  • Query complexity
  • Prometheus cardinality
  • Large time ranges
  • Data source latency

23. Grafana + OpenShift Monitoring

Since you’ve worked with OpenShift, expect these:

What components are included?
  • Prometheus
  • Alertmanager
  • Grafana (user-managed)
  • Thanos
  • Node Exporter
  • kube-state-metrics
How do you monitor OpenShift?

Monitor:

  • Nodes
  • Pods
  • Namespaces
  • etcd
  • API Server
  • OVN Networking
  • Ingress Controllers

24. Scenario Question

Interviewer: Dashboard loading takes 30 seconds. What do you check?

Answer:

  1. Query Inspector
  2. Prometheus query duration
  3. Cardinality explosion
  4. Recording Rules
  5. Dashboard refresh interval
  6. Data source latency
  7. Browser performance

25. Senior-Level Question

How would you design observability for 500 Kubernetes/OpenShift nodes?

Answer:

Node Exporter
kube-state-metrics
OpenTelemetry
Prometheus HA
Mimir
Grafana
Alertmanager

Include:

  • Long-term retention in Mimir
  • Loki for logs
  • Tempo for tracing
  • HA Prometheus
  • Multi-cluster federation
  • SSO integration
  • RBAC
  • Disaster Recovery

Interview Questions Often Asked for Senior Grafana Architects
  1. How do you reduce Prometheus cardinality?
  2. Difference between recording rules and alert rules?
  3. How do you monitor Kubernetes at scale?
  4. How would you deploy Grafana HA?
  5. How would you design LGTM for 1000 servers?
  6. How do you troubleshoot missing metrics?
  7. How do you secure Grafana in an enterprise environment?
  8. How do you integrate Grafana with OpenShift?
  9. How do you implement multi-tenancy?
  10. How do you monitor AWS, Azure, and OpenShift from a single Grafana instance?

These are the questions typically asked for Senior Monitoring Engineer, SRE, OpenShift Architect, and Cloud Architect interviews.

Leave a Reply