Site Reliability Engineer

Aethir β€’ Taiwan

Company

Aethir

Location

Taiwan

Type

Full Time

Job Description

Aethir is the only Enterprise-grade AI-focused GPU-as-a-service provider in the market. Its decentralized cloud computing infrastructure allows GPU providers (containers) to meet Enterprise clients who need powerful GPU chips for professional AI/ML tasks. Thanks to a constantly growing network of over 40,000 top-shelf GPUs, including 3,000 NVIDIA H100s, Aethir is able to provide enterprise-grade GPU computing wherever it’s needed, at scale.

Backed by leading Web3 investors like Framework Ventures, Merit Circle, Hashkey, Animoca Brands, Sanctor Capital, Infinity Ventures Crypto (IVC), and others, with over $130M in funds raised for the ecosystem, Aethir is paving the way for the future of decentralized computing.

We are looking for an operations and maintenance development engineer (SRE) to join our new headquarters in Kuala Lumpur, Malaysia, who will play a critical role in monitoring, troubleshooting, and optimizing our production system to ensure the highest levels of performance and stability for our AI and gaming customers worldwide.

Responsibilities
  • Monitor, Review, and Respond to Faults: Take on the responsibility of monitoring, reviewing, responding to faults, troubleshooting, resolving, and subsequently optimizing the production system.
  • System Architecture and Performance: Continuously monitor and review the system architecture, process logic, system performance, stability, and other technical areas and indicators to ensure their rationality.
  • Coordination with Business Team: Drive the business team in resolving any issues related to operations and maintenance.
  • Production Failure Response: Respond promptly to production failures, acting as the overall coordinator for resolution.
  • Collaborative Problem-Solving: Organize relevant R&D, operations and maintenance, and product teams to collaboratively investigate and resolve problems.
  • Failure Response Time: Responsible for the failure response time and resolution time, ensuring timely resolution of issues.
  • Case Studies and Optimization: Conduct case studies on production issues and follow up with optimizations to improve system performance and stability.
  • Documentation: Maintain comprehensive documentation of system architecture, processes, and troubleshooting procedures.
  • Continuous Improvement: Identify areas for improvement in the operations and maintenance processes and implement necessary changes.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related field.
  • Experience in operations and maintenance development, preferably in a cloud computing or AI-focused environment.
  • Strong understanding of system architecture, performance monitoring, and troubleshooting methodologies.
  • Excellent communication and collaboration skills.
  • Ability to work in a fast-paced, startup environment.
  • Proficiency in Kubernetes (K8S), CI/CD, and Docker.
  • Expertise in AWS (VPC, S3, EC2, etc.) or Python (one of the two).
  • Responsible for building the operations and maintenance infrastructure platform and handling core business operations.
  • Management experience is a plus, but not required.
  • Prior experience working in structured environments such as Huawei, ZTE, or banking institutions is preferred.

Benefits

  • Hypergrowth Startup Environment
  • Fantastic Career Progression Opportunities
  • Work within a Global and Local Team
  • Collaborative and innovative work environment with opportunities to contribute to cutting-edge projects.


About the company

Aethir is revolutionizing DePIN with its advanced, distributed enterprise-grade GPU-based compute infrastructure tailored for AI and gaming. Backed by leading Web3 investors like Framework Ventures, Merit Circle, Hashkey, Animoca Brands, Sanctor Capital, Infinity Ventures Crypto (IVC), and others, with over $140M in funds raised for the ecosystem, Aethir is paving the way for the future of decentralized computing.

Apply Now

Date Posted

10/16/2024

Views

0

Back to Job Listings ❀️Add To Job List Company Info View Company Reviews
Positive
Subjectivity Score: 0.8

Similar Jobs

Junior Engineer - Hitachi Energy

Views in the last 30 days - 0

The job posting is for an Engineering position with a focus on projects of low to medium complexity The role involves providing technical support coll...

View Details

Product Engineer - Eaton

Views in the last 30 days - 0

The Taiwan Product Engineer will lead product development in the Tainan Taiwan manufacturing facility coordinating crossfunctional teams for quick sam...

View Details

Software Engineer III, Machine Learning, Camera - Google

Views in the last 30 days - 0

Googles Devices Services team is seeking a software engineer with experience in software development data structures and algorithms The ideal candida...

View Details

Software Engineer II - Cadence

Views in the last 30 days - 0

Cadence is seeking leaders and innovators with advanced degrees in Computer Science or Electrical Engineering Key responsibilities include developing ...

View Details

Principal Solutions Engineer - Cadence

Views in the last 30 days - 0

Cadence is seeking an experienced programmer to develop highperformance PERC decks The role involves working with toptalented customers and internal t...

View Details

Automation PC-Base Engineer - Jabil

Views in the last 30 days - 0

Jabil a global leader in engineering manufacturing and supply chain solutions is seeking an Automation Design Engineer II The role involves designing ...

View Details