
1001 - 5000 employees
Founded 2008
💼 Consulting
🏥 Healthcare
📦 Logistics
💰 $4.1M Venture Round on 2013-01
Consulting • Healthcare • Logistics
Cloudera is a leading enterprise data cloud company that empowers businesses to manage and analyze data across any environment. Offering a hybrid data platform, Cloudera facilitates modern data architectures with solutions like open data lakehouse, scalable data mesh, and unified data fabric, designed for artificial intelligence, data engineering, and machine learning. Key industries served include financial services, telecommunications, healthcare, and more, where Cloudera's platform enables secure, scalable, and effective data management. By leveraging AI and advanced analytics at scale, Cloudera helps organizations transform their data into actionable insights.
🔥 0 minutes ago
🌐 Spain, Hungary, +1 more countries – Remote
⏰ Full Time
🟡 Mid-level
🟠 Senior
👷🏻♀️ Engineer
👻 Ghost score 10%
Apache
AWS
Cloud
Docker
Google Cloud Platform
Grafana
Kafka
Kubernetes
Microservices
Neo4j
Postgres
SDLC
SQL
Terraform
Vault
Improve your chances of getting an interview by checking your resume score before you apply.

1001 - 5000 employees
Founded 2008
💼 Consulting
🏥 Healthcare
📦 Logistics
💰 $4.1M Venture Round on 2013-01
Consulting • Healthcare • Logistics
Cloudera is a leading enterprise data cloud company that empowers businesses to manage and analyze data across any environment. Offering a hybrid data platform, Cloudera facilitates modern data architectures with solutions like open data lakehouse, scalable data mesh, and unified data fabric, designed for artificial intelligence, data engineering, and machine learning. Key industries served include financial services, telecommunications, healthcare, and more, where Cloudera's platform enables secure, scalable, and effective data management. By leveraging AI and advanced analytics at scale, Cloudera helps organizations transform their data into actionable insights.
• Provision, tune, and maintain production-grade Neo4j graph database and pgvector vector storage clusters • Engineer high-throughput index structures, cosine similarity vector indexes, and query optimizations for sub-second responses • Build automated ingestion pipelines parsing Git repositories, ASTs, Jira issue links, Apache Avro schemas, and CI/CD metadata into an enterprise knowledge graph • Connect distributed pipeline engines to hybrid retrievers combining SQL, Cypher graph traversals, and dense vector embeddings • Configure circuit breakers, confidence scoring thresholds, and step-limit constraints for autonomous agents • Integrate microservices and knowledge stores with the Enterprise AI Gateway • Maintain version-controlled system prompt structures in localized .ai/ spoke directories while following DLP PII scrubbing rules and token rate limits • Implement automated failover, backup restoration, and multi-cloud storage-tier cost controls across AWS and GCP • Own the semantic, vector, and graph storage layer powering the context engine for enterprise AI utilities and the Internal Developer Portal • Lead deployment of the SDLC Context Graph and GraphRAG Engine for CAB compliance, code/schema lineage tracking, and enterprise LLM proxy integrations
• Deep operational and development experience with Neo4j (Cypher, APOC, causal clustering) or enterprise Knowledge Graphs • Proven expertise with pgvector (PostgreSQL), embeddings management, hybrid search techniques, and framework integrations (LangChain, LlamaIndex, or custom RAG pipelines) • Hands-on experience managing relational (PostgreSQL) and graph databases across AWS and GCP cloud environments • Proficiency in consuming Apache Avro payloads, streaming Kafka events (AWS MSK), and parsing structured/unstructured code and JSON artifacts • Practical understanding of Prompts-as-Code patterns, few-shot prompt optimization, and agent tool specification • Experience with Infrastructure-as-Code (Terraform) primitives, Kubernetes (EKS/GKE), Docker, and pull-based GitOps workflows • Exposure to HashiCorp Vault Transit encryption, OIDC keyless authentication, and zero-trust workload identities • Familiarity with OpenTelemetry (OTel) instrumentation for tracking vector search query latencies and LLM inference performance in Datadog or Grafana
• Generous PTO Policy • Unplugged Days supporting work-life balance • Flexible WFH Policy • Mental & Physical Wellness programs • Phone and Internet Reimbursement program • Access to Continued Career Development • Comprehensive Benefits and Competitive Packages • Paid Volunteer Time • Employee Resource Groups
Apply Now🕒 6 days ago
Ingeniero/a Safety desarrollando análisis de riesgos, documentación CENELEC y Safety Cases para sistemas CBTC ferroviarios en Expleo. Colaboración con equipos de señalización, validación y certificación.
🗣️🇪🇸 Spanish Required
🕒 September 22
Bioinformatics Engineer building Nextflow workflows and cloud infrastructure for Seqera’s scientific data platform. Supporting customers with reproducible pipelines for genomics, imaging, and large-scale scientific datasets.
AWS
Azure
Cloud
Groovy
Python
🕒 September 22
Senior Knowledge Graph Engineer building GraphRAG, entity-resolution, and semantic data systems. Powering EcoVadis sustainability ratings and autonomous AI agents from Barcelona or remotely within Spain.
Azure
Cloud
ETL
Neo4j
Python
🕒 September 21
Site and Commissioning Engineer installing and commissioning Kardex AutoStore systems at customer sites. Managing site execution, technical testing, customer training, and go-live support across EMEA.
🕒 September 21
Spares Engineer identificando piezas obsoletas y evaluando riesgos para ATEXIS, consultora multinacional de ingeniería aeroespacial. Proponiendo soluciones alternativas con experiencia en materiales compuestos y cadena de suministro.
🗣️🇪🇸 Spanish Required