Chess

Site Reliability Engineer

Remote
5+ years exp
Full-time
Posted 2d ago
0 views
Actively Hiring Direct 1-Click Apply

Check Your Resume Match Score

Scan your resume against ATS criteria for this Site Reliability Engineer role at Chess.

Apply for this position

Apply on Company Website

Job Description

About the role

The Site Reliability Engineer will play a critical role in ensuring the stability, performance, and scalability of our global gaming platform infrastructure. This position exists to bridge the gap between development and operations, maintaining high availability for millions of concurrent users while supporting rapid feature development and deployment. The SRE will be instrumental in building resilient systems that can handle massive scale across multiple regions, directly impacting user experience and platform reliability.

As our platform continues to grow and serve a global community, this role will drive the technical infrastructure decisions that enable seamless gaming experiences. The position requires both deep technical expertise and collaborative leadership to work across engineering teams, ensuring our systems can scale efficiently while maintaining the performance standards our users expect.

What you'll do

Design and implement

multi-regional resilient infrastructure capable of handling millions of concurrent sessions and transactions daily across global data centers

Lead

the hybrid cloud migration strategy, integrating bare-metal datacenter resources with cloud services for optimal performance and cost efficiency

Own

the on-call rotation and incident response procedures, ensuring rapid resolution of critical system issues and maintaining high availability SLAs

Architect

monitoring and alerting systems using industry-standard tools to proactively identify and resolve performance bottlenecks before they impact users

Collaborate

with development teams to implement infrastructure-as-code practices and establish deployment pipelines that support continuous integration and delivery

Optimize

system performance through capacity planning, load testing, and resource allocation across distributed computing environments

Establish

and maintain security protocols and risk assessment procedures for infrastructure components and data protection

Partner

with engineering teams to design scalable solutions for high-traffic applications and real-time processing requirements

Drive

automation initiatives to reduce manual operational overhead and improve system reliability through scripting and configuration management

Mentor

team members on SRE best practices and contribute to the development of infrastructure standards and documentation

Preferred Skills

  • Bachelor's degree in Computer Science, Engineering, or related technical field, or equivalent practical experience
  • 5+ years of experience in site reliability engineering, DevOps, or infrastructure engineering roles
  • Experience managing bare-metal server infrastructure and datacenter operations
  • Strong proficiency with UNIX/Linux operating systems and command-line administration
  • Experience with cloud platforms (GCP, AWS, or Azure) and infrastructure-as-code tools (Terraform, CloudFormation, or similar)
  • Hands-on experience with configuration management systems (Ansible, Chef, Puppet, or similar)
  • Solid understanding of networking fundamentals, protocols (TCP/IP, HTTP/HTTPS, DNS), and network troubleshooting
  • Experience with containerization and orchestration technologies (Docker, Kubernetes, or similar)
  • Proficiency with monitoring and observability tools (Datadog, Prometheus, Grafana, ELK stack, or similar)
  • Experience with relational and NoSQL databases, including performance optimization and scaling strategies
  • Strong collaboration and communication skills for working effectively in a distributed team environment
  • Demonstrated sense of ownership and accountability for system reliability and performance

Nice to have

  • Advanced knowledge of content delivery networks (CDNs) and edge computing
  • Experience with server-side automation and scripting languages (Python, Go, Bash, or similar)
  • Background in high-availability architectures and disaster recovery planning
  • Familiarity with security frameworks and compliance requirements
  • Experience with game server infrastructure or real-time application hosting
  • Knowledge of database administration and optimization for high-concurrency applications
  • Understanding of CI/CD pipelines and deployment automation
  • Experience with capacity planning and performance testing tools
  • Previous experience in a fully remote, distributed work environment
  • Continuous learning mindset with interest in emerging infrastructure technologies

About the Opportunity

  • This is a full-time opportunity
  • We are 100% remote (work from anywhere!)

You can learn more about us here:

Key Requirements & Skills

  • Advanced knowledge of content delivery networks (CDNs) and edge computing
  • Experience with server-side automation and scripting languages (Python, Go, Bash, or similar)
  • Background in high-availability architectures and disaster recovery planning
  • Familiarity with security frameworks and compliance requirements
  • Experience with game server infrastructure or real-time application hosting
  • Knowledge of database administration and optimization for high-concurrency applications
  • Understanding of CI/CD pipelines and deployment automation
  • Experience with capacity planning and performance testing tools
  • Previous experience in a fully remote, distributed work environment
  • Continuous learning mindset with interest in emerging infrastructure technologies

Frequently Asked Questions

How to apply for Site Reliability Engineer at Chess?

Click the "Apply via CareerScan" button on this page.

What is the salary for this role?

Salary details will be discussed during the interview.

What experience is required?

5+ years of experience is required.

Is this position still open?

Yes, currently active and accepting applications.

Site Reliability Engineer

Chess · Remote