(USA) Staff, Software Engineer
WalmartAbout the role
Position Summary...
What you'll do...
About the Role:
We are looking for a Staff Software Engineer who is an expert in Python, with a deep mastery of LLM-driven systems, agentic architectures, and scalable intelligent automation. This is a pivotal role where you'll define and implement the technical strategy, influence architecture decisions, and lead the creation of next-gen AI-driven platform that move beyond prompt engineering into modular reasoning system.
This role isn’t about building one-off workflows—it’s about inventing and hardening intelligent systems that can reason, act, and adapt. You will shape the core architecture of multi-agent platforms, ensure LLM integrations are secure, efficient, and observable, and build frameworks that others can extend across use cases and orgs.
Key Responsibilities:
Architect modular, testable, and composable Python systems that support multi-agent workflows, tool-chaining, RAG, memory management, and fallback strategies.
Design LLM-powered execution engines that support both high throughput and adaptive reasoning (via LangChain, AutoGen, or custom frameworks).
Lead implementation of retrieval-augmented generation (RAG) pipelines, semantic search, and structured knowledge memory systems.
Build and scale integrations with internal LLMs, including handling signature-based auth, function calling, and context management at scale.
Drive end-to-end lifecycle: from configuration schema (YAML) to execution trace logging, observability, and self-healing recovery patterns.
Must-Have Qualifications:
8 + years of professional software engineering experience, with 5+ years in Python, building distributed systems at scale.
Deep knowledge of agentic design patterns, including:
ReAct, Plan-and-Execute, AutoGen-style coordination
Tool calling, dynamic agent routing, and recursive agent planning
Semantic memory, embedding-based context lookup, summarization windows
Expertise in building LLM-based systems with LangChain, OpenAI, Anthropic, or custom orchestrators.
Hands-on experience with:
RAG pipelines using vector stores (FAISS, Pinecone, Weaviate, Qdrant, Azure Cognitive Search)
LLM evaluation and observability (tracing, token usage, agent state tracking)
Workflow orchestration using config-first approaches (YAML/JSON definitions, step runners, etc.)
Proven ability to drive technical vision, resolve ambiguity, and make architectural tradeoffs at scale.
Strong background in distributed systems, task queues, asynchronous workflows, and backend performance optimization.
Experience in cloud-native environments (AWS, GCP, or Azure), including containerization, monitoring, and secure API integrations.
Nice-to-Have:
Built or contributed to a custom agentic orchestration framework used across multiple product lines.
Experience with vector search optimization, context ranking, or temporal memory solutions.
Published talks, blogs, or papers on LLM systems, AI architecture, or applied reasoning frameworks.
Deep understanding of how to apply LLM systems in regulated or high-compliance environments (PII handling, redaction, observability).
Exposure to DevEx platforms for developers to build workflows on top of intelligent agents.
Familiarity with multi-modal agents (text + vision), LLM simulation patterns, or offline evaluation loops.
What Success Looks Like:
You’ve built an agent platform that others across the org use as the foundation for intelligent automation.
You turn abstract ideas into clean, extensible, production-grade Python systems that scale and evolve.
You elevate technical conversations, coach Staff+ engineers, and are the go-to person for unblocking complex challenges.
You are not just LLM-aware—you pioneer how LLMs are applied in production systems, with a clear perspective on wh
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s