Principal Software Engineer

16 hours ago

fort myers, florida, United States Velastegui Ventures Full-time

Senior Software Engineer

  Full-Time · Remote (U.S.)


  About Velastegui Ventures

  We build enterprise AI infrastructure that lets regulated organizations query their unstructured and structured data in natural language. Our platform combines retrieval-augmented generation, federated search, and self-hosted model  inference to deliver citation-backed, auditable answers — deployed inside the customer's own cloud boundary. We're growing the engineering team to scale this platform through its first enterprise deployments and into a multi-tenant product.


  About the Role

  You'll design, build, and operate the backend of our Enterprise Knowledge and Natural Language Query Platform: ingestion, retrieval, inference, and the governance layer that makes AI defensible for regulated customers. This is a hands-on  senior role with real architectural ownership — you'll ship on the platform as it stands today while helping design the next generation of it. We're looking for someone comfortable operating anywhere from 5 to 15 years into their career, provided the depth is there.


  Where the platform is today, and where it's going

  Today, ingestion runs on SQS with Ray-Serve workers, retrieval uses Amazon Bedrock and Bedrock Knowledge Bases inside the customer's AWS boundary, and the platform is deployed via a GitOps supply chain (ArgoCD, Cosign, Kyverno) into customer-owned EKS.


  Over the next 12–18 months we're evolving toward a Kafka-based streaming backbone with CDC, self-hosted LLM inference (vLLM / SGLang), self-hosted vector and graph retrieval, and federated search across semantic, lexical, and  knowledge-graph modalities. You'll work on both sides of that transition.


  The first 6 months

  You'll help stand up our first BYOC deployment into an enterprise client on the current SQS + Bedrock stack, own one or two of the six data-source connectors end-to-end, and drive the ingestion pipeline's incremental-sync design.

  Expect direct contact with the customer's platform and compliance teams. In parallel, you'll help shape the Kafka-based connector framework and self-hosted inference path that lands on the platform later in the year.


  What You'll Do

  - Design and operate backend services across an event-driven microservices architecture — ingestion, retrieval, inference, and the connector framework that onboards new data sources.

  - Ship on the current SQS + Ray-Serve + Bedrock stack while helping design and land the Kafka + CDC + self-hosted-inference generation of it.

  - Extend retrieval quality across semantic, lexical, and (eventually) graph modalities, including fusion and re-ranking.

  - Grow the AI serving path: embedding, generation, re-ranking, model routing, and the registries and evals that let us promote models safely.

  - Build observability and evaluation into every service you own — tracing, metrics, and quality/drift signals, not just system health.

  - Partner with security and compliance on the controls that make SOC 2 and (where applicable) HIPAA and GDPR defensible: encryption, access control, audit lineage, data residency.

  - Support enterprise deployments alongside customer platform teams — IaC rollout, cutover, rollback, and the day-two handoff.


  Required

  - 5+ years building and operating production backend systems, with meaningful ownership of distributed architecture.

  - Strong proficiency in at least one of Go, Rust, or Python, and comfort working across a polyglot codebase.

  - Solid grounding in distributed systems: asynchronous messaging, event-driven design, and the failure modes that come with them. Direct experience with a streaming or queue technology (Kafka, SQS, RabbitMQ, or similar) and the concepts

  behind Change Data Capture.

  - Working knowledge of PostgreSQL and at least one non-relational store (vector, document, or graph).

  - Fluency with AWS (or equivalent), containers (Docker/Kubernetes), and infrastructure-as-code (Terraform or similar).

  - Real interest in — ideally hands-on experience with — applied AI/ML systems: embeddings, vector search, RAG, or LLM serving. This is an AI infrastructure role; curiosity here isn't optional.

  - Communication skills to work directly with customers, compliance stakeholders, and product.


  Nice to Have (any of these move you up the stack)

  - Managed AI services and hosted retrieval (Amazon Bedrock, Bedrock Knowledge Bases, OpenAI / Anthropic APIs) — relevant to the current stack.

  - Self-hosted LLM serving (vLLM, SGLang, llama.cpp) or OpenAI-compatible gateways — relevant to where we're going.

  - Vector databases and hybrid semantic/lexical