English summary for screening — check the original posting before applying.
Lead reliability design for all systems of "Mecha Comic" (24 million monthly users). This role involves defining reliability standards across the entire platform, from architecture selection to security, and improving the overall operational level through technical guidance. You will work in a large-scale environment utilizing both AWS and Google Cloud.
Must-haves
- 3+ years of experience in infrastructure design, environment setup, and operations/maintenance for web systems.
- Understanding of standard TCP/IP networking and common protocols like DNS and HTTP.
Nice-to-haves
- Experience defining SLO/SLI and implementing error budget operations.
- Experience driving reliability improvement projects across multiple systems.
- Experience with architecture design, environment setup, and operations/maintenance of cloud services (AWS/GCP).
- Experience with performance tuning and troubleshooting for high-load web systems.
- Experience in web application development.
- Experience with security management for web systems and CSIRT/PSIRT activities.
- Experience designing and improving on-call systems and incident response processes.
- Experience establishing SRE teams or mentoring members.
Tech stack
AWSGoogle CloudEC2ECSAuroraElastiCacheDynamoDBLambdaWAFS3OpenSearchEMRGlueFirehoseBigQueryGCSVertex AINewRelicDatadogCloudWatchCloudFormationTerraformGitGitHub EnterpriseSlackRedmineGoogle MeetZoomTypeScriptRubyReactRuby on RailsPostgreSQLFaaSCaaSCloud RunCloud FunctionsCloud BatchObject StorageXDR
Work style
Hybrid (2 days remote per week), Tokyo, Japan
Other notes
Annual salary range: 8,000,000 - 12,000,000 JPY