Skip to content
View onkar717's full-sized avatar
💥
💥

Block or report onkar717

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
onkar717/README.md

Onkar Shelke

Site Reliability Engineer · Kubernetes · Cloud native · Open source

I keep production Kubernetes fleets healthy and fix the open source tools they run on.

Open to SRE and Platform roles LinkedIn Email dev.to X

$ kubectl get engineer onkar-shelke -o wide. NAME onkar-shelke, ROLE SRE, CLUSTERS 50+, MERGED-PRS 120+, CERTS CKA, KCNA, AWS-SAA, RHCSA, STATUS Available.

About · Experience · Open source · Talks and mentoring · Projects · Writing · Certifications · Stack

👋 About me

I'm an SRE based in Pune, India. At Obmondo I ran a GitOps fleet of 50+ production Kubernetes clusters across AWS, Azure, Hetzner and bare metal, and owned the storage underneath it. At the same time I was a maintainer of KubeStellar, a CNCF Sandbox project, where I also mentored LFX contributors. In 2026 I completed Google Summer of Code with the JBoss Community, contributing to jws-diag, a diagnostic CLI for Red Hat's JBoss Web Server.

  • 🔧 Now: contributing upstream to Kubernetes SIG projects (Kueue, kind, Kubebuilder), Hive and jws-diag
  • ✍️ Writing: debugging stories from Kubernetes and Prometheus on dev.to
  • 🤝 Open to: SRE, Platform and Infrastructure roles. Remote in any time zone, or on site in Pune or Bangalore.
🧭 Quick facts
Roles I'm targeting Site Reliability Engineer · Platform Engineer · Infrastructure Engineer
Experience SRE at Obmondo · KubeStellar maintainer and LFX mentor · CNCF LFX intern · GSoC 2026
Production scale 50+ Kubernetes clusters · 150+ nodes · 8+ enterprise tenants · 524.8 TB of ZFS storage
Open source Merged upstream PRs PRs in review
Certifications CKA · KCNA · AWS Solutions Architect Associate · RHCSA
Speaking KCD Delhi 2026 · Cloud Native Pune
Education B.E. Computer Engineering, PDEA College of Engineering, Pune · 2026 · CGPA 8.46
Location Pune, India (IST, UTC+5:30)

⚙️ Production experience

Site Reliability Engineer · Obmondo · remote · 2025 to 2026

  • Ran a multi-cloud GitOps fleet of 50+ production Kubernetes clusters on AWS, Azure, Hetzner and bare metal for 8+ enterprise tenants, shipping 30+ Helm applications through ArgoCD and Puppet across 150+ nodes.
  • Owned the storage stack: Ceph (25+ runbooks, OSD recovery in rescue mode, live Docker to containerd migration), ZFS RAID-Z3 pools holding 524.8 TB on 48-disk configs, and Harbor (database moved to CloudNativePG, garbage collection freeing 194 GB+ per cycle).
  • Led the RCA for a platform outage caused by two failures at once: Cilium config drift (a missing --devices flag) and a ZFS LocalPV DaemonSet scheduling mistake that evicted a Redis StatefulSet. Service was back in 3 hours through ArgoCD reconciliation and node label fixes.

🌍 Open source

The counts below are live. Click any number to see the pull requests behind it.

Project My role Merged PRs What I worked on
KubeStellar
CNCF Sandbox · multi-cluster Kubernetes
Maintainer (UI approver, 2025 to 2026) and LFX mentor Multi-cluster topology views, in-browser pod logs and exec, Helm chart deployment tab. Reviewed 50+ contributor PRs and onboarded 6+ new contributors.
jws-diag
Red Hat JBoss Web Server (Tomcat)
GSoC 2026 contributor summary, config, logs, diff and instances commands, a server.xml parser that applies Tomcat defaults, TLS parsing with credential redaction, and an 80% test coverage gate.
Kueue
Kubernetes SIG Scheduling
Contributor Fixed silently dropped errors in SparkApplication volumes. Tests for the kueueviz backend.
Apache Tomcat Contributor Unit tests for six valves, including AccessLogValve, PersistentValve and SemaphoreValve.
Hive
AI agent orchestration for OSS maintainers
Contributor Self-hosted Kubernetes fixes: RBAC steps, storage and probe docs, image pull policy.
OpenEverest
CNCF Sandbox · databases on Kubernetes
Contributor Event replay buffer so reconnecting subscribers don't lose events.
Apache SkyWalking BanyanDB Contributor Dynamic reload of TLS certificates and keys.

✅ Latest merged

🔄 In review

Selected pull requests, grouped by the kind of work (click to expand)

Reliability fixes

  • openeverest#2591: replay buffer so subscribers that reconnect don't lose events
  • kueue#13548: stop silently dropping errors in SparkApplication addVolumes
  • skywalking-banyandb#642: reload TLS certificates and keys without a restart
  • hive#9334: always pull the floating image in the self-hosted Deployment
  • jws-diag#65: read listening sockets from /proc instead of binding ports
  • jws-diag#67: check the Tomcat process user, not the user running the tool

Features

Tests and quality

  • jws-diag#27: JaCoCo reporting with an 80% instruction coverage floor
  • tomcat#965, #976, #979: tests for JsonErrorReportValve, FilterValve, ProxyErrorReportValve, SemaphoreValve, AccessLogValve and PersistentValve
  • wildfly-elytron#2422: tests for OidcClientUriBuilder
  • kueue#13488 and #13550: kueueviz handler tests, a fixed sleep replaced with a stability poll

Community

  • cncf/mentoring#1464 and #1472: LFX mentorship projects for KubeStellar UI (Developer Relations, Model Context Protocol)

🎤 Talks and mentoring

  • KCD Delhi 2026 (21 Feb 2026): AI-Assisted Kubernetes Policy Management with Kyverno, with Akshay Kumar. Agenda
  • Cloud Native Pune (CNCG x Docker): KubeStellar for multi-cluster and edge workload placement
  • LFX Mentorship, CNCF: mentor on 4 KubeStellar UI projects in 2025, including the Developer Relations and Community Growth track
  • GeeksforGeeks Campus Mantri: ran coding workshops and contests at my college

🛠️ Projects

VisualEyes veye CLI showing system metrics, Kubernetes health and a firing alert

Open source observability platform for Kubernetes with an AI-assisted incident response engine.

  • Go multi-module backend with system and Kubernetes agents, Prometheus metrics and a 12-command Bubbletea CLI
  • 6-agent RCA pipeline (CrewAI) that streams progress over SSE and classifies incidents SEV1 to SEV4
  • 7 guarded kubectl remediation tools behind a shell-injection allowlist, with Slack alerts on SEV1 and SEV2

Go · Python · CrewAI · Kubernetes · Prometheus · PostgreSQL · SQLite

✍️ Writing

📜 Certifications

CKA badge
CKA
Certified Kubernetes Administrator
Verify on Credly
KCNA badge
KCNA
Kubernetes and Cloud Native Associate
Verify on Credly
AWS Certified Solutions Architect Associate badge
AWS SAA
Solutions Architect Associate
SAA-C03
Red Hat Certified System Administrator badge
RHCSA
Red Hat Certified System Administrator
EX200

🧰 Stack

Kubernetes, Docker, AWS, Azure, GCP, Terraform, Linux, Red Hat
Prometheus, Grafana, PostgreSQL, Redis, MongoDB, MySQL, GitHub Actions, Jenkins
Go, Java, Python, TypeScript, JavaScript, C++, Bash, Git

Area Tools
Kubernetes Kubernetes, Cluster API, Helm, ArgoCD, Cilium, Kyverno, Velero
Storage and data Rook-Ceph, ZFS, Harbor, CloudNativePG, PostgreSQL, Redis, MongoDB, MySQL, OpenSearch
Observability Prometheus, Grafana, Alertmanager, incident response and RCA
Cloud and IaC AWS, Azure, GCP, Hetzner, bare metal, Terraform, Puppet
CI/CD GitHub Actions, Jenkins, SonarQube, Docker
Languages Go, Java, Python, TypeScript, JavaScript, C/C++, Bash

📈 Contribution activity

Contribution graph animation

Hiring for SRE or platform work? Email onkarwork2234@gmail.com or message me on LinkedIn.

The merged PRs, PRs in review, posts and contribution graph refresh every day through a GitHub Actions workflow in this repo.

Popular repositories Loading

  1. istio istio Public

    Forked from istio/istio

    Connect, secure, control, and observe services.

    Go 1

  2. VisualEyes VisualEyes Public

    Open-source cloud-native observability platform system & Kubernetes monitoring, AI-powered RCA, alerting, and live CLI dashboard

    Go 1

  3. Assignment-2 Assignment-2 Public

    Lets Upgrade Assignment-2

    HTML

  4. CodSoft CodSoft Public

    HTML

  5. Coders-Cave-Phase1 Coders-Cave-Phase1 Public

    Phase 1

    HTML

  6. Bhart-Intern Bhart-Intern Public

    Onkar Raghinath Shelke

    HTML