Data Engineer (AI Agents)
Sigmatic
- Ubicación
- Remoto
- Salario
- USD 3,000 – 5,000
- Publicada
- Hoy
- Fuente
- Get on Board
What we are looking for Required • Five or more years in data engineering or analytics engineering, having owned pipelines a business depended on • Strong SQL and Python: production transformations, testing, debugging, performance work • Serious dbt experience. Multi-tenant or multi-source projects are what we most want to hear about • ETL and ELT against messy real sources: APIs, databases, files, cloud storage, with full and incremental loads, CDC and idempotent reruns • A way of working that already assumes AI assistance, and a clear account of how you verify what you did not write • Precision in writing. Much of your output is read by a model before a person sees it, so unambiguous descriptions of what a column means are a deliverable, not an afterthought • A record of finishing: edge cases, validation, documentation, deployment, production support
How we build We would rather be straight about the difficulty than sell you a tidy version of it. How we build Most of our code is now written with AI assistance, including substantial autonomous work. The scarce skill on this team is no longer producing code. It is specifying work precisely enough to delegate, and verifying output rigorously enough to trust. In practice that means we ablate tests rather than admiring them: delete the guard, confirm the test goes red. A test that still passes with the code removed is worse than no test. It means a bug fix starts by proving the test fails against the unfixed code. And it means that when we do not know how a vendor behaves, we probe it and write down the number instead of reasoning about it. You do not need experience with any particular tool. You do need to be comfortable in a codebase where much of the diff was not typed by a human, and to have real opinions about how to verify it. If that is already how you work you will move fast here. If it is not, this will be a frustrating role. Our stack • Transformation — dbt on PostgreSQL: staging, intermediate and marts, with contracts and tests enforced in CI • Warehouse — multi-tenant PostgreSQL on AWS RDS, schema per tenant • Extraction and loading — Python on AWS Lambda, Glue PySpark, Step Functions, EventBridge, with SAM and Terraform • Sources — EHR and clinical systems, accounting (QuickBooks, NetSuite), scheduling, supply chain, IoT and scanner event streams • AI layer — AWS Bedrock, a generated semantic catalogue, Qdrant retrieval, Python agent services • Engineering — GitHub with gitflow, mandatory review, CI gates, Linear, SOC 2 Type II
Strong plus • AWS data stack with operational ownership: Glue, Lambda, Step Functions, EventBridge • Multi-tenant SaaS platforms where customer-specific schemas map into a shared model • Layered architecture with write-audit-publish or equivalent quality gating • Data products consumed by LLMs or agents: semantic layers, catalogues, retrieval Useful • Healthcare operational, revenue-cycle, claims, scheduling or EHR-adjacent data; HIPAA-aware architecture, de-identification, least privilege • Accounting and financial system data, operational KPI development, IoT or time-series
