| Position Details | |
| Job Title | Data Engineer |
| Company | Santech Business Solutions LLC |
| Location | Katy, TX |
| Interview | Video |
| Salary | Based on Experience (BOE) |
| Job Type | Full Time · On-Site |
| Req ID | BPD-2026-DE-001 |
About the Company
Santech America is an early-stage Industrial IoT (IIoT) and Industrial AI startup building next-generation asset performance monitoring and predictive maintenance solutions for manufacturing companies. Our platform bridges physical equipment and actionable intelligence — from sensor-level data acquisition on the plant floor to cloud-hosted dashboards and ML-driven maintenance recommendations.
We operate a lean, high-ownership engineering team under Santech Business Solutions LLC, headquartered in Katy, TX. If you thrive in a zero-to-one environment and want your work to directly impact production uptime at manufacturing plants, this role is built for you.
About This Opportunity
We are looking for a Data Engineer to own and scale our end-to-end data pipeline — from edge telemetry ingestion (Modbus TCP, InfluxDB at the edge) to multi-cloud data infrastructure, Agentic AI pipelines, and ML-ready feature engineering. You will work at the intersection of time-series sensor data, multi-cloud platforms (Azure, AWS, GCP), advanced Python data engineering, and machine learning operations.
This is a full-time, on-site position based in Katy, TX. You will work directly with the founding engineering team with high ownership and a clear path to Lead / Principal Data Engineer as the company scales.
Job Description
Data Pipeline & Ingestion
- • Design, build, and maintain high-reliability data pipelines that move time-series sensor telemetry from edge hardware to cloud-hosted InfluxDB instances
- • Manage and optimize Telegraf dual-output configuration for edge-to-cloud InfluxDB replication, ensuring low-latency, fault-tolerant data flow
- • Develop and maintain event-driven pipelines using Azure Service Bus (or equivalent) for discrete operational events — alerts, ticket triggers, and sensor state changes
Multi-Cloud Infrastructure
- • Architect and manage data pipelines across Microsoft Azure (Service Bus, Azure ML, Azure DevOps), AWS (S3, Glue, Kinesis, Lambda, SageMaker), and GCP (BigQuery, Pub/Sub, Cloud Dataflow, Vertex AI)
- • Support customer-specific deployment patterns including data isolation to customer-owned cloud subscriptions across multiple CSPs using Terraform and IaC
- • Design cloud-agnostic data storage and ingestion layers enabling portability between Azure, AWS, and GCP without major rearchitecting
- • Integrate ERP systems (e.g., Oracle Fusion) to support MTO/BTO parts replacement workflows triggered by predictive maintenance signals
Agentic AI & Advanced Python
- • Build and integrate Agentic AI pipelines — autonomous multi-agent workflows that reason over sensor data, trigger maintenance actions, and generate operational insights using frameworks such as LangChain, LangGraph, AutoGen, or CrewAI
- • Develop data pipelines using advanced Python libraries: pandas, NumPy, Polars, PySpark for large-scale data processing; SQLAlchemy, PyArrow for data access and quality; Pydantic, FastAPI for data contract and API layers
- • Design tool-calling and retrieval-augmented generation (RAG) patterns for AI agents that query structured sensor databases and unstructured maintenance logs
- • Support MLflow and Azure ML / SageMaker / Vertex AI experiment tracking with data versioning, lineage tracking, and preprocessing reproducibility
- • Apply SMOTE and missing data imputation techniques for handling imbalanced and incomplete industrial datasets
Storage & Database Engineering
- • Design, administer, and tune a wide range of databases: PostgreSQL, Oracle, MySQL, Microsoft SQL Server (relational); MongoDB, Cosmos DB (document/NoSQL); InfluxDB, TimescaleDB (time-series); Redis (caching); and cloud-native warehouses such as BigQuery, Redshift, Synapse Analytics
- • Build data models supporting installation tracking, device registry (sensor-to-gateway mapping), and tenant-level data isolation across multi-cloud environments
- • Design schemas, retention policies, and downsampling strategies for high-frequency sensor time-series data
- • Implement cross-database ETL and CDC (Change Data Capture) patterns for real-time and batch data synchronization
Data Quality & Observability
- • Implement data quality checks, schema validation, and alerting for pipeline anomalies (sensor dropouts, Modbus polling failures, replication lag)
- • Build observability dashboards in Grafana to monitor pipeline health, sensor uptime, and ingestion throughput
- • Establish data governance standards including audit trails for installation records and sensor provenance across cloud environments
Required Qualifications
- • 2+ years of data engineering experience with demonstrated ownership of production data pipelines
- • Advanced proficiency in Python — pandas, NumPy, Polars, PySpark, Dask, SQLAlchemy, Great Expectations, Pydantic, FastAPI
- • Hands-on experience with multi-cloud platforms: Microsoft Azure, AWS, and/or GCP — data services, compute, and storage layers
- • Production experience with relational databases: PostgreSQL, Oracle, MySQL, and/or SQL Server — schema design, query optimization, indexing
- • Experience with NoSQL / document databases: MongoDB, Cosmos DB, or equivalent
- • Hands-on experience with InfluxDB or equivalent time-series databases (TimescaleDB, ClickHouse)
- • Experience building Agentic AI or LLM-orchestration workflows using LangChain, LangGraph, AutoGen, CrewAI, or similar frameworks
- • Familiarity with Docker and container-based deployments (Docker Compose, CI/CD integration)
- • Understanding of message queue and event-driven architectures (Azure Service Bus, AWS Kinesis / SQS, GCP Pub/Sub, or Kafka)
- • PLC Software Knowledge — familiarity with PLC programming environments (e.g., Siemens TIA Portal, Rockwell Studio 5000 / RSLogix, Beckhoff TwinCAT) and understanding of ladder logic, function block diagrams (FBD), or structured text (ST) for industrial automation context
- • Experience with Grafana or comparable monitoring and visualization platforms
- • Solid SQL skills for data modeling, analytical queries, and ETL logic
Nice to Have
- • Prior background in Industrial IoT or manufacturing data environments — OT/IT convergence, SCADA, Modbus, OPC-UA
- • Knowledge of Telegraf for metrics collection and routing
- • Experience with MLflow or cloud-native MLOps pipeline management (Azure ML, SageMaker, Vertex AI)
- • Infrastructure-as-code experience using Terraform or ARM/Bicep/CloudFormation templates
- • Exposure to ERP integration patterns (Oracle Fusion, SAP) for enterprise data workflows
- • Familiarity with SMOTE, class imbalance handling, and advanced data preprocessing for ML
Education
- • Bachelor’s Degree Required — Computer Science, Data Science, Electrical Engineering, or related field preferred
- • Master’s Degree Preferred — focus in Data Science, Industrial Engineering, or Systems Engineering is a plus
- • Equivalent scope of experience in a production data engineering role will be considered in lieu of a degree
How to Apply
Submit your resume and a brief note on your relevant data engineering experience — specifically around multi-cloud pipelines, database engineering, Agentic AI, or industrial/IoT environments — to hr@santechamerica.com.
Applicants must have current and unrestricted authorization to work in the United States. This position is not able to provide visa sponsorship at this time. Future sponsorship eligibility will be evaluated based on business needs and is not guaranteed.
Santech Business Solutions LLC is an Equal Opportunity Employer. We do not discriminate based on race, color, religion, sex, national origin, age, disability, veteran status, or any other characteristic protected by applicable law.