Big Data Infrastructure Engineer

Job not on LinkedIn

🔥 0 minutes ago

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of LigaData

LigaData

51 - 200 employees

🤖 Artificial Intelligence

📡 Telecommunications

🤝 B2B

Artificial Intelligence • Telecommunications • B2B

LigaData is a company that provides data and AI-driven telecom products and services tailored for communications service providers. Their offerings include the LigaData Telecom Data Fabric, which unifies and correlates subscriber and operations data, and the LigaData Data Platform, which offers innovative and cost-effective ways to manage and analyze large volumes of data. They also offer AI apps to address telecom data and analytics use cases, and provide various managed services and professional services to support their products. LigaData partners with major cloud providers like Google Cloud Platform, Microsoft Azure, and AWS to enhance their telecom tech solutions with AI and cloud innovations.

📋 Description

• Leverage AI-assisted tools to improve troubleshooting, log analysis, scripting, documentation, research, and operational efficiency while validating outputs before implementation • Identify repetitive operational activities and develop automation solutions using Shell/Bash, Python, Ansible, APIs, or other appropriate technologies • Evaluate emerging AI and automation capabilities and identify opportunities to improve infrastructure operations and engineering workflows • Administer, configure, manage, troubleshoot, and optimize Linux operating systems in production and non-production environments • Monitor and analyze CPU, memory, disk, filesystem, network, processes, and system services; perform configuration and performance tuning • Administer PostgreSQL, MySQL/MariaDB, and Redis, including configuration, access management, backup and recovery, monitoring, troubleshooting, maintenance, and performance optimization • Support database replication, high availability, backup/recovery, and capacity management • Support and maintain Docker and Kubernetes environments, including deployment, configuration, monitoring, troubleshooting, scaling, and cluster administration • Support clustered and distributed platforms, focusing on high availability, replication, failover, load balancing, quorum, capacity management, and disaster recovery • Support installation, configuration, monitoring, administration, and upgrades of Cloudera/Hortonworks and Hadoop-based environments • Support and troubleshoot HDFS, YARN, Hive, Spark, HBase, Kafka, Airflow, Superset, and Trino/Presto • Perform production monitoring and support using Zabbix and Grafana; participate in incident management, root cause analysis, and corrective/preventive actions • Support security integrations and technologies such as Ranger, LDAP, and Kerberos • Collaborate with development, infrastructure, and other technical teams on deployments, upgrades, infrastructure changes, troubleshooting, and production support • Maintain technical documentation, operational procedures, automation, and infrastructure configuration records

🎯 Requirements

• 3–5 years of relevant hands-on experience in Linux/System Administration, Database Administration, Big Data Infrastructure, DevOps, or a related infrastructure role • Bachelor's Degree in Computer Science, Computer Engineering, Information Technology, or a related field • Strong AI-first and automation-driven mindset, with demonstrated ability to use AI-assisted tools effectively in technical workflows and critically validate generated recommendations before applying them • Good scripting and automation skills using Shell/Bash; knowledge of Python, Ansible, APIs, or similar technologies is highly desirable • Strong hands-on knowledge of Linux administration, including system configuration, service management, resource management, storage/filesystems, permissions, networking, troubleshooting, and performance optimization • Good hands-on knowledge of PostgreSQL, MySQL/MariaDB, and Redis administration, including configuration, backup and recovery, users and privileges, monitoring, maintenance, and performance tuning • Good understanding of database concepts including connections, transactions, locks, indexing, query performance, replication, and high availability • Good hands-on understanding of Docker and Kubernetes, including containers, images, pods, deployments, services, storage, networking, monitoring, resource management, and troubleshooting • Good understanding of clustering and distributed system concepts, including high availability, replication, failover, load balancing, and quorum • Good understanding of networking fundamentals, including TCP/IP, DNS, ports, routing, connectivity, and network troubleshooting • Good understanding of Big Data concepts and the Hadoop ecosystem, with familiarity or hands-on experience in HDFS, YARN, Hive, Spark, Kafka, and HBase • Familiarity with Cloudera or Hortonworks platforms is highly desirable • Familiarity with Zabbix/Grafana, Ranger/LDAP/Kerberos, CI/CD tools, and Trino/Presto is an advantage • Strong troubleshooting, analytical, and problem-solving skills, with the ability to investigate issues systematically and identify root causes • Ability to work effectively in production environments, collaborate across technical teams, take ownership of assigned activities, and continuously develop technical knowledge • Preferred certifications include RHCSA, RHCE, CKA, PostgreSQL or MySQL-related certifications/training, and Red Hat Ansible or other relevant automation certifications; certifications are advantageous and not a substitute for practical hands-on experience

🏖️ Benefits

• Remote work arrangement • Full-time employment

Apply Now

Similar Jobs

🔥 1 hour ago

All Native Group, The Federal Services Division of Ho-Chunk Inc.

501 - 1000

🎖️ Defense

🏥 Healthcare

📦 Logistics

Data Architect managing WRAIR technology-transfer data, agreements, licensing, and research partnerships. Maintaining databases, CRADA invoicing, intellectual property workflows, and researcher support.

🔥 1 hour ago

Wynd Labs

11 - 50

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

Data Engineer building web-scale crawling, scraping, and dataset pipelines for companies training powerful AI models. Operating distributed systems for videos, transcripts, audio, and public web data.

🔥 3 hours ago

Gravie

201 - 500

🏥 Healthcare

🛡️ Insurance

🤝 B2B

Senior Data Engineer building and operating Gravie’s near-real-time AWS streaming platform. Securing PHI and transforming operational data for health-benefits products.

🇺🇸 United States – Remote

💵 $133.9k - $178.5k / year

💰 $150M Private Equity Round - Gravie on 2025-05

⏰ Full Time

🟠 Senior

🚰 Data Engineer

🔥 4 hours ago

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Data Engineer leading Databricks schema, security, governance, and migration work. Integrating federal marine and terrestrial species data for Peraton’s government program.

🔥 4 hours ago

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Cloud Data Architect leading Databricks lakehouse and ETL architecture for Peraton’s federal ESA species-data integration program. Designing schema harmonization, metadata, and data-quality frameworks in hybrid AWS cloud.