An A-people software development company that provides high-quality outsourcing services to the US and Europe. We work with the best people giving them the ability to work with interesting projects and modern technologies. We strive to build relationships with our clients that last for years.
15 липня 2026

Senior Data Engineer (GCP, PySpark, Dataproc) (вакансія неактивна)

віддалено

We are looking for a Senior Data Engineer to join Implex and work on a long-term data platform project for a public-sector in the Middle East.

The role combines hands-on PySpark development with the configuration and troubleshooting of data integrations across GCP Dataproc, S3-compatible storage, PostgreSQL, and SAP HANA.

You will work with large-scale ETL pipelines, investigate Spark and infrastructure-level issues, and help ensure reliable and secure data exchange between multiple systems.

Responsibilities

  • Develop, maintain, and optimize ETL pipelines using Apache Spark and PySpark.
  • Configure and troubleshoot Spark workloads running on GCP Dataproc.
  • Package and submit jobs, manage dependencies, and analyze driver and executor logs.
  • Identify and resolve Spark performance, memory, partitioning, and data processing issues.
  • Integrate Spark with S3-compatible storage using the Hadoop S3A connector.
  • Configure and troubleshoot MinIO connectivity, bucket access, TLS endpoints, signatures, and redirects.
  • Configure custom CA certificates and Java truststores for Spark drivers and executors.
  • Implement scalable Spark JDBC reads and writes for PostgreSQL and SAP HANA.
  • Build reliable incremental data loads, retries, backfills, and data reconciliation processes.
  • Contribute to data modeling, schema evolution, documentation, and data quality practices.

Must-Have Technical Skills

  • Apache Spark / PySpark development (Dataproc): driver/executor behavior, job packaging/submission, performance tuning
  • GCP Dataproc operations: cluster configuration, init actions, dependency management, troubleshooting via logs/metrics
  • Hadoop S3A connector: `fs.s3a.*` configuration, endpoint/path-style access, credential providers, S3 semantics
  • MinIO (S3-compatible) integration: bucket policies, TLS endpoints, signature/redirect troubleshooting
  • TLS/SSL & PKI with custom CA: certificate chains, SAN/hostname validation, diagnosing handshake/PKIX errors
  • Java truststores (JKS/PKCS12) & JVM SSL config: `keytool`, distributing truststores, setting driver/executor JVM options
  • PostgreSQL integration: Spark JDBC reads/writes at scale, indexing/performance basics, data type mapping
  • SAP HANA integration: JDBC/ODBC connectivity, driver management, calculation views vs tables, pushdown/performance tuning
  • ETL engineering: incremental loads/CDC concepts, idempotency, retries, backfills, data quality/reconciliation
  • Data Warehousing integration: strong SQL, staging-to-publish patterns, SCD concepts, bulk load strategies
  • Data modeling & governance basics: dimensional modeling, schema evolution, lineage/documentation practices

Nice-to-Have Skills (but not must)

  • Linux + networking fundamentals: DNS, routing, firewall/LB/proxy basics; tools like `curl`/`openssl s_client` for validation
  • Secure secrets handling: GCP Secret Manager (or equivalent), least-privilege access, avoiding hardcoded credentials
  • KAFKA knowledge if we ever bring KAFKA into the architecture again

Hiring process

  • HR interview
  • Technical interview with Implex
  • PM interview on the client side
  • Dev Lead interview on the client side

What we offer

  • A long-term international project
  • Opportunity to work on a national-scale digital platform used by thousands of users
  • Remote full-time collaboration
  • Professional and supportive team environment
  • Challenging technical tasks and a meaningful product with real-world impact