Service Record · Form CS-02 · Rev. MMXXVI

Character Name Andrew Caylor

Class
Senior Platform & Reliability Engineer
Level
99
Experience
7+ years
Certification
CKAD

§ IProfessional Summary

Senior Platform and Reliability Engineer with 7+ years of experience designing and operating large-scale distributed infrastructure across AWS, GCP, Azure, and on-prem environments. Specialized in Kubernetes platform engineering, observability systems, GitOps delivery, and infrastructure automation supporting mission-critical workloads.

Download as PDF

§ IISkills

Cloud & Platforms

  • Multi-cloud architecture across AWS, GCP, and Azure.
  • Kubernetes platform engineering with EKS, GKE, RKE2, OpenShift, and AKS.
  • Infrastructure automation with Terraform, Python, Crossplane, and Ansible.

CI/CD & GitOps

  • GitHub Actions, GitLab CI, Jenkins, Argo CD, Helm, and Kustomize.
  • Continuous delivery platforms for high-change production environments.

Telemetry & Reliability

  • Prometheus, Grafana, Loki, Mimir, InfluxDB, OpenTelemetry, and Alertmanager.
  • Incident command, on-call operations, proactive alerting, and dashboard design.

Backend & Systems

  • Python, JavaScript/TypeScript, PostgreSQL, Elasticsearch, Kafka, and Redis.
  • Linux, networking, DNS, PKI, TLS, object storage, and secrets management.

AI Engineering

  • AI coding agents including Claude Code, Codex, Cursor, and Opencode.
  • Self-hosted AI models, MCP servers, coding agent plugins, and AI-assisted PR reviews.

§ IIICampaign History

Redox Inc.

Senior Software Engineer — Platform

  • Designed and operated multi-cloud Kubernetes infrastructure in AWS and GCP supporting 75+ production services and reducing deployment time by 15%.
  • Designed and implemented a GitOps-based continuous delivery platform using Crossplane and Argo CD, reducing deployment time by 80% across production services.
  • Led implementation of a centralized observability and telemetry platform using Prometheus, Grafana, and OpenTelemetry, enabling proactive alerting for 75+ production services and delivering 20+ multi-cloud, multi-environment dashboards.
  • Served as primary on-call engineer and incident commander for production incidents and maintenance events, improving system reliability and reducing operational response times.
  • Authored 15+ architecture proposals and drove cross-team platform initiatives impacting platform reliability, deployment workflows, and observability practices.
  • Built internal tooling and automation leveraging LLM-based workflows to reduce dependency management and platform maintenance effort.

Georgia Tech Research Institute

Research Faculty — Cloud Engineer

  • Built standardized Kubernetes-based application platforms supporting research workloads and production services.
  • Led multi-cloud infrastructure initiatives in Amazon Web Services and Microsoft Azure.
  • Led development of a scalable remote learning lab platform used by 20+ students per class across 3 curricula.
  • Designed and deployed a high-availability on-prem Kubernetes cluster integrated with enterprise identity systems and VMware infrastructure.

Georgia Tech Research Institute

Research Faculty — Systems Administrator

  • Automated large-scale bare-metal and VM deployments using Foreman, Red Hat Linux, Python, Packer, and Ansible.
  • Designed and maintained enterprise infrastructure spanning virtualization, storage fabrics, networking, and operating systems.
  • Implemented centralized log aggregation and observability tooling, reducing time to query system logs from 5 minutes to under 30 seconds.
  • Led migration of critical file storage infrastructure to a distributed SAN architecture, improving uptime to 99.99%.
  • Developed disaster recovery and business continuity strategies reducing MTTR to under one day.

Georgia Tech Research Institute

IT Internship

  • Managed VMware vSphere virtualization stack, including storage, networking, and hypervisor infrastructure across three data centers with 1 PB of highly available SAN storage.
  • Led IT process automation projects, customizing virtual machines and streamlining software deployment and upgrades.
  • Resolved escalated IT helpdesk tickets for Active Directory, Windows Server, Red Hat Enterprise Linux, DHCP, DNS, and organization-wide CIFS/NFS file sharing.

§ IVNotable Feats

Multi-Cloud Elasticsearch Platform

  • Led design and implementation of a Kubernetes-native platform using the ECK operator to provide a repeatable, cloud-agnostic document database across AWS, GCP, and Azure.
  • Architected a highly available cluster supporting 100 TB+ of indexed data and serving 2,000+ application pods as a primary backend.
  • Improved query latency by about 20% and throughput by about 10% while standardizing monitoring and alerting with Prometheus and Terraform.

High-Scale Metrics & Observability Platform

  • Led re-architecture of a single-node metrics pipeline into a distributed Kubernetes-native platform using InfluxDB, Telegraf, and Grafana Mimir.
  • Designed and implemented a Helm-based deployment with S3-backed storage, replacing custom cloud-init workflows and enabling consistent installs across environments.
  • Scaled the system to support 42+ million active metric series in InfluxDB and 12 million series in Mimir.
  • Reduced InfluxDB cold boot time from about 40 minutes to under 5 minutes and enabled zero-downtime patching for Mimir.

§ VCredentials

Professional Certifications

  • Certified Kubernetes Application Developer (CKAD)

Education

Kennesaw State University, Georgia
Bachelor of Business Administration in Information Systems
May 2019