Site Reliability Engineer, Ark Large Model Platform

TikTok Singapore

Company

TikTok

Location

Singapore

Type

Full Time

Job Description

Responsibilities

TikTok will be prioritizing applicants who have a current right to work in Singapore, and do not require TikTok's sponsorship of a visa.

TikTok is the leading destination for short-form mobile video. Our mission is to inspire creativity and bring joy. TikTok has global offices including Los Angeles, New York, London, Paris, Berlin, Dubai, Singapore, Jakarta, Seoul and Tokyo.

Why Join Us
Creation is the core of TikTok's purpose. Our platform is built to help imaginations thrive. This is doubly true of the teams that make TikTok possible.
Together, we inspire creativity and bring joy - a mission we all believe in and aim towards achieving every day.
To us, every challenge, no matter how difficult, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always.

Want more jobs like this?

Get Software Engineering jobs in Singapore delivered to your inbox every week.

By signing up, you agree to our Terms of Service & Privacy Policy.

At TikTok, we create together and grow together. That's how we drive impact - for ourselves, our company, and the communities we serve.
Join us.

About the Team
The Applied Machine Learning (AML) - Enterprise team provides machine learning platform products on VolcanoEngine with cloud native resource scheduling system which intelligently orchestrates different tasks and jobs with minimised costs of every experiment and maximised resource utilisation, rich modelling tools including customised machine learning tasks and web IDE, and multi-framework high performance model inference services.

In 2021, through VolcanoEngine, we released this machine learning infrastructure to the public, to provide more enterprises with reduced costs of computation power, lower barriers to machine learning engineering and deeper developments in AI capabilities.

Responsibilities
Responsible for Ark Large Model Platform development on Volcano Engine, researching systematic solutions on large model solution implementations and applications in various industries, striving to reduce the IT cost of large model applications, meeting the users' ever-growing demand for intelligent interaction and improving the lifestyle and communications of users in the future world.

- Manage and oversee the stability of both control and data aspects of large-scale model systems through effective DevOps practices.
- Develop and enhance observability systems for monitoring the stability of large model systems, ensuring high reliability and performance.
- Handle super large-scale cluster management and ensure efficient operation and maintenance of large model systems.

Qualifications

- B. Sc or higher degree in Computer Science or related fields from accredited and reputable institutions with R&D experience in the fields of cloud computing or large-scale model systems.
- Proficiency in cloud-native technologies and understanding of the relevant technology stack.
- Expertise in one of the following programming languages: Golang, Python, or Java, with the ability to use it proficiently in a professional setting.
- Familiarity with cloud-native technologies for log collection, monitoring, and alerting.

Preferred Qualifications:
- Prior experience in the construction and maintenance of stability systems for large-scale infrastructures.
- Experience in operating and maintaining large-scale systems.
- Experience with infrastructure as code, particularly Terraform, is highly desirable.

TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Apply Now

Date Posted

01/23/2025

Views

0

Back to Job Listings ❤️Add To Job List Company Info View Company Reviews
Positive
Subjectivity Score: 0.8

Similar Jobs

Sales Engineer - MariaDB plc

Views in the last 30 days - 0

MariaDB is a leading database for modern application development used by 75 of the Fortune 500 and billions of people daily The company is seeking a S...

View Details

Partner Manager - MariaDB plc

Views in the last 30 days - 0

MariaDB is a leading database for modern application development used by 75 of the Fortune 500 and billions of people daily The Partner Manager role i...

View Details

Revenue Systems Manager (HubSpot Admin) - UpGuard

Views in the last 30 days - 0

UpGuard is seeking a Revenue Systems Manager to manage and improve their tech stack ensuring seamless integration across systems The role involves sys...

View Details

Regional Vice President - Sales - Graylog, Inc

Views in the last 30 days - 0

Graylog a renowned centralized log management and Security Information Event Management SIEM provider is seeking a Regional Vice President of Sales fo...

View Details

Machine Learning Engineer (Recommendation), TikTok Global e-Commerce - 2025 Start - TikTok

Views in the last 30 days - 0

TikTok is seeking applicants with the right to work in Singapore for a role in their Ecommerce team The team focuses on applied machine learning and d...

View Details

Test Engineer III - Jabil

Views in the last 30 days - 0

Jabil a global leader in engineering manufacturing and supply chain solutions is seeking Direct Test Engineers I II The role involves designing devel...

View Details