ZT
open to devops & sre roles

~/whoami

Zokchen Tamang

DevOps Engineer II — Kathmandu, Nepal

I build infrastructure that doesn't wake me up at 3am. Off the clock you'll find me chasing a personal best on a run or grabbing boards on the basketball court — on the job, I bring the same drive to keeping systems alive under pressure. Linux administration and networking are where I'm strongest, backed by hands-on Kubernetes, GitOps, and cloud work.

containers managed
500+
uptime
99.9%
servers managed
15+
services on gitops
8
aws stack cost
$30/mo

Reliability isn't a checkbox — it's the win condition.

I'm a sole-charge DevOps engineer operating production infrastructure end to end: Linux and networking, containers and orchestration, CI/CD, observability, and security — across 500+ containers running at once, spanning databases, applications, and everything in between for the enterprise, including a network of 25+ college websites alongside our core IAM and CLMS platforms. I deploy and maintain a wide range of application stacks (Django, Laravel, Node.js, React, Angular, Java, Odoo) backed by PostgreSQL, MySQL, and SQL Server, all watched over by an observability stack I built myself with Grafana, Prometheus, Loki, and the Elastic Stack.

I stood up a K3s cluster from scratch, layered in HashiCorp Vault for secrets and ArgoCD for GitOps, and I like passing that knowledge on — I've mentored a junior engineer to the point he can run the infrastructure solo when I'm on leave.

The stack I run in production

linux_networking

Linux server administration, UFW firewalls, DNS, SSL/TLS, Cloudflare, reverse proxy routing, performance tuning

observability

Grafana, Prometheus, Loki, Promtail, Elastic Stack (Elasticsearch & Kibana 8.x/9.x)

cloud_aws

EC2, S3, Route 53, EBS, EFS — architected and hosted a client subscription service at ~$30/month

containers

Docker 23.0.5, Docker Compose v2.17.3, Kubernetes (K3s v1.32.6, Kustomize v5.5.0), Longhorn, Portainer, Komodo

gitops_secrets

ArgoCD managing 8 services via monorepo, webhook-triggered deployments, custom Helm charts, Traefik DNS routing; HashiCorp Vault (KV engine) with per-developer access control

ci_cd

GitLab CI, GitHub Actions, Harbor (private registry with image scanning & vulnerability checks), automated build / test / deploy

scripting

Bash scripting for Loki log parsing by service, server hardening scripts, 25+ cron jobs for automated backups (rclone → cloud storage), Docker pruning, Elasticsearch snapshots

databases

PostgreSQL (15, 16, 18), MySQL, SQL Server (Windows Server) across DEV/INT/PROD

app_stacks

Python/Django (Celery, Redis, RabbitMQ, Flower), Node.js, PHP/Laravel, Java, React, Angular, Odoo (v16, v18)

version_control

Git, GitLab, GitHub

Where I've kept things running

2024

present

DevOps Engineer II

Innovate Nepal Group · Kathmandu, Nepal

  • Sole DevOps engineer operating and monitoring 15+ Linux servers across DEV, INT, and PROD, sustaining 99.9% uptime, secured with UFW firewalls and Fail2Ban hardening.
  • Built and operate a 3-node K3s cluster on VMs (1 master, 2 workers) with Longhorn for persistent storage; created custom Helm charts with Traefik handling DNS-based domain routing.
  • Run ArgoCD GitOps across 8 services (5 Django microservice stacks, 3 Java stacks) deployed from a single monorepo via webhook-triggered pipelines.
  • Implemented HashiCorp Vault for centralized secrets with per-developer, per-environment access control and automatic app restarts on secret changes.
  • Built a full observability stack (Grafana, Prometheus, Loki, Promtail) plus an Elastic Stack for the Java application, cutting mean time to detect incidents by an estimated 40%.
  • Hosted a private Harbor container registry with image scanning and vulnerability checks, integrated into CI/CD.
  • Automated backups with 25+ cron jobs, syncing to cloud storage via rclone — cutting manual backup time from 2-3 days to a scheduled job.
  • Discovered a critical JavaScript injection vulnerability during an infrastructure review; remediated it with a Caddy WAF and rate limiting.
  • Trained and mentored a junior DevOps engineer across Docker, K3s, and networking — he's now capable of covering full infrastructure operations when I'm on leave.

2023

present

Freelance DevOps Engineer

Self-Employed · Remote (part-time)

  • Architected and hosted a client subscription service on AWS (EC2, S3, Route 53, EBS, EFS) with NGINX Proxy Manager, a fully Dockerized application, and a GitHub Actions pipeline — optimized to run at roughly $30/month.
  • Delivered end-to-end DevOps for multiple clients: Linux server setup, Dockerized deployments, and CI/CD via GitHub Actions.
  • Managed DNS, SSL certificates, and reverse proxies using Cloudflare and NGINX Proxy Manager.

2021

2023

Earlier roles

Data Analyst · Cloud Factory (Nov 2021 – Mar 2022)

Data Manager & Customer Service Rep. · Best Himalaya Pvt. Ltd. (Aug 2019 – Aug 2023)

A few systems I've shipped

JavaGitLab CIHarborK3sGrafana

Enrollment & Attendance Management System

Problem
A 15+ microservice Java application was deployed manually through Jenkins, with no registry and no way to deploy individual services.
Approach
Built a custom GitLab CI pipeline with conditional logic for per-service and full-stack releases, tuned with optimized Docker images, and stood up a private Harbor registry integrated into the pipeline.
Outcome
Pipeline time cut from 10–15 minutes to 3–5 minutes, fully automated from a developer's push.
Python/DjangoDocker ComposeGitLab CI

Multi-Tenant Django Platform

Problem
Scaling from 3 to 8 tenants exposed unbounded CPU/memory growth from Django's default worker scaling, consuming nearly 100% of server resources.
Approach
Diagnosed the resource pattern and applied container-level CPU/memory limits; fine-tuned the monorepo GitLab CI pipeline.
Outcome
Freed ~12GB of system memory (10–15% of total) and cut CI time from 7–8 minutes to 2–3 minutes.
PHP/LaravelDockerPrivate Registry25+ sites

IAM, CLMS & the ING College Network

Problem
IAM and CLMS are core platforms I run in production; alongside them I maintain deployments for the full ING college network — 25+ websites in total, including Islington College and Herald College Kathmandu — all running unoptimized Docker images at real production scale.
Approach
Standardized and optimized the Docker image build process across all 25+ sites plus IAM and CLMS, rolling out one consistent private-registry pipeline enterprise-wide instead of per-site one-offs.
Outcome
Cut image sizes by ~80% and sustained 99.07% uptime across the network; CLMS alone handles 300–400 concurrent users at peak, while the college sites are SEO-optimized for high-traffic public access.
Node.jsDockerGitHub ActionsCaddy

DajuVai.com E-Commerce Platform

Problem
Every release required manually SSHing in and running a git pull, taking several minutes per deploy with no rollback path.
Approach
Replaced it with a GitHub Actions pipeline, multi-stage Docker builds for smaller images, and a self-hosted production environment with Caddy auto-SSL.
Outcome
Deploys now ship in 10–20 seconds, fully automated from a push.

Let's build something reliable.

Open to DevOps, Cloud, and SRE roles — happy to talk about infrastructure, GitOps, or your next production incident over email.