Avance Consulting

Site Reliability Engineer

Die Stellenanzeige deiner ursprünglichen Anfrage ist leider nicht mehr vakant.
Wie wäre es mit diesem Job-Angebot?

Position Overview

  • Security Clearance & Vetting Level: Public Sector Clearance + NdK (Nachweis der Kundigkeit)

The Observability & SRE Engineer builds and operates central telemetry stacks to provide visibility across distributed cloud infrastructure. This role implements metric collection, log aggregation, distributed tracing, and standalone alerting tailored for Air-Gap operations.

Key Responsibilities

  • Monitoring Stack Setup: Deploy and maintain Prometheus, Grafana, Loki, and Jaeger stacks as code.
  • Dashboards & Alerting: Develop PromQL/LogQL monitoring dashboards; configure AlertManager routing and inhibition for autonomous Air-Gap operations.
  • Distributed Tracing: Implement OpenTelemetry Collectors and SDKs for auto-instrumentation and OTLP trace propagation.
  • SIEM Integration: Configure secure log export and security event management (CEF/Syslog) for external SIEM platforms.

Technical Qualifications & Skills

  • Must Have:
  • Deep expertise in Prometheus (PromQL, ServiceMonitor, Federation, Remote Write) and Grafana.
  • Hands-on experience with Loki log aggregation and AlertManager routing.
  • Proficiency with OpenTelemetry (Collectors, SDKs, OTLP) and Jaeger distributed tracing.
  • Knowledge of SIEM integrations and security event logging.
  • Scripting skills in Go, Python, Shell, and YAML.

Weitere interessante Stellenanzeigen