Abstracts

Applications to speak. These go through review and a decision.

TagsSpeakerRatingsFiles
Docs That Answer Back: Retrieval-Grounded Documentation SitesPendingApprovedAbstractA 10-minute tour of turning a static docs site into one that answers questions with citations, stays honest when it doesn't know, and costs under $50/month to run. Live demo, real failure cases, and a checklist you can apply to your own docs this week.Sswyx+kms-0033-speaker@ai.engineerSswyx+kms-0033-speaker@ai.engineer
Lightning: Agents in Production Q&AAcceptedApprovedYesAbstractA rapid-fire Q&A lightning talk on running AI agents safely and reliably in production, covering observability, guardrails, and rollback strategies.Ssbek-speaker2@example.comSsbek-speaker2@example.com
Your AI Pair Programmer Is Lying to You: Verification Patterns That ScalePendingApprovedAbstractCode generation is easy; trusting it is hard. This session covers verification patterns for AI-generated code — property tests, mutation coverage, snapshot judges, and CI gates — with data from 18 months of running them on a 200-engineer codebase. Includes what we stopped doing because it didn't catch anything.Ssbek-speaker2@example.comSsbek-speaker2@example.com
Docs That Answer Back: Retrieval-Grounded Documentation SitesAccept QueueApprovedAbstractA 10-minute tour of turning a static docs site into one that answers questions with citations, stays honest when it doesn't know, and costs under $50/month to run. Live demo, real failure cases, and a checklist you can apply to your own docs this week.AOAda OkonkwoAOAda Okonkwo
Taming 40-Minute CI: Incremental Builds at Monorepo ScaleAcceptedApprovedYesAbstractOur monorepo CI took 40 minutes on a good day. This talk walks through how we cut it to 6 minutes with content-addressed caching, remote execution, and a test-selection model — including the two migrations that failed first. You'll leave with a decision framework for which incremental-build investments pay off at which repo sizes, and the graphs to convince your platform team. Updated: now includes 2026 benchmark data.AOAda OkonkwoAOAda Okonkwo3.63 (2)
Your AI Pair Programmer Is Lying to You: Verification Patterns That ScaleAccept QueueApprovedAbstractCode generation is easy; trusting it is hard. This session covers verification patterns for AI-generated code — property tests, mutation coverage, snapshot judges, and CI gates — with data from 18 months of running them on a 200-engineer codebase. Includes what we stopped doing because it didn't catch anything.AOAda OkonkwoAOAda Okonkwo5.00 (1)
Full Journey Demo Talk (Round 2)AcceptedApprovedYesAbstractA clean re-run of the full submit-to-publish journey, with the Organiser assigning the abstract to a committee, a Reviewer blind-scoring it, and the Organiser deciding, notifying, scheduling and publishing based on that score.Ffull-journey-demo-2@example.comFfull-journey-demo-2@example.comOct 11, 2026, 12:00 AMRoom B4.50 (1 of 4)
Reviewer-First Journey TalkAcceptedApprovedYesAbstractA demo submission created to walk the proper submit-review-accept-schedule-publish journey, with the reviewer step included this time.Rreviewer-first-demo@example.comRreviewer-first-demo@example.comOct 10, 2026, 11:00 PMRoom B4.00 (1 of 4)
Full Journey Walkthrough TalkAcceptedApprovedYesAbstractA live demo walkthrough talk created to show the full submit-to-publish journey end to end.Wwalkthrough-demo@example.comWwalkthrough-demo@example.comOct 10, 2026, 10:00 PMRoom A
Claude - EventPendingIn ReviewAbstractTestFfarishussain021@gmail.comFfarishussain021@gmail.com
E2E Journey Test TalkAcceptedApprovedYesAbstractA test submission created to walk the full submit-to-publish journey end to end.Ee2e-tester@example.comEe2e-tester@example.comOct 10, 2026, 8:00 PMRoom A
Closing RemarksPendingApprovedAbstractAI-039Closing thoughts, thanks, and what we would like to see submitted next year.Talk (30 min)AdvancedEnglishEvaluationOpen SourceSecuritySRSofia RossiSOSam Organiser0.9
Opening Keynote: Why NowPendingApprovedAbstractAI-040Why this conference, why now, and what we hope you take away from the next two days.Talk (30 min)BeginnerEnglishProductOpen SourceWCWei ChenSOSam Organiser1.2
Panel: The Ethics of AutonomyPendingApprovedAbstractAI-038Autonomy raises questions the industry has been deferring. A panel on accountability, disclosure and the decisions we are quietly making on users' behalf.Talk (30 min)IntermediateEnglishInfrastructureCostSecurityIKIdris KhanSOSam Organiser0.2
Panel: Buy, Build or WaitDraftApprovedAbstractAI-037Three practitioners argue about when to buy, when to build and when to wait. Audience questions throughout; the panel has been asked not to agree.Talk (30 min)BeginnerEnglishApplied AIAgentsCostNFNaomi FischerWCWei ChenSOSam Organiser0.1
Workshop: Production RAG in Half a DayWithdrawnApprovedAbstractAI-035Half a day taking a RAG prototype to something you would put in front of customers: chunking, retrieval quality, evaluation and the deployment checklist.Talk (30 min)IntermediateEnglishEvaluationOpen SourceRAGPRPriya RamanSOSam Organiser2.1
Lightning: Five Prompts That Changed Our ProductDraftApprovedAbstractAI-036Five prompts that measurably changed a product metric, what they replaced, and why each one worked. Five minutes, no slides after the first.Workshop (120 min)AdvancedEnglishProductAgentsRAGTBTomas BergSOSam Organiser1.4
Workshop: Building Your First EvaluatorWithdrawnApprovedAbstractAI-034A hands-on session building an evaluator from scratch. Bring a laptop; you will leave with a working harness scoring your own outputs against a labelled set you build in the room.Talk (30 min)BeginnerEnglishInfrastructureOpen SourceSecurityLMLuis MoreauSOSam Organiser1.3
Debugging Agent LoopsDeclinedApprovedYesAbstractAI-030Agent loops fail in ways stack traces do not capture. The trace format we adopted, how we replay a failing run locally, and the visualisation that made the loops legible to people who did not write them.Talk (30 min)AdvancedEnglishInfrastructureOpen SourceRAGSWSam WhitfieldSOSam Organiser0.4
Batching, Queueing and BackpressureDeclinedApprovedYesAbstractAI-028Queue depth, batch size and tail latency pull against each other. What we learned running a shared inference tier, and the backpressure design that stopped one team's spike from taking down everyone.Talk (30 min)BeginnerEnglishProductCostKMKofi MensahSOSam Organiser0.5
Migrating Off a Single VendorDeclinedApprovedYesAbstractAI-029Moving off a single provider without a rewrite. The compatibility layer, the differences that leaked through anyway, and an honest cost comparison after the dust settled.Workshop (120 min)IntermediateEnglishApplied AISecurityEPElena PetrovaADAmara DialloSOSam Organiser1
Semantic Caching in PracticeDecline QueueApprovedAbstractAI-026Caching on meaning rather than exact match, and the ways that goes wrong. Threshold tuning, the confidently-wrong cache hit, and how we detect when the cache is degrading answer quality.Talk (30 min)IntermediateEnglishInfrastructureRAGJHJonas HalvorsenSOSam Organiser4.50 (1 of 4)1.5
The Hidden Cost of Context WindowsDecline QueueApprovedAbstractAI-027Long context is not free and not always better. Measured cost and quality across context lengths on a real workload, plus the retrieval strategy that beat simply pasting everything in.Talk (30 min)AdvancedEnglishEvaluationAgentsMCMira CastellanosSOSam Organiser2.4
Open Weights in the EnterpriseDecline QueueApprovedAbstractAI-025Open-weight models inside a company with a compliance function. Licensing, hosting, evaluation against hosted alternatives, and the security questions that decided it for us.Talk (30 min)BeginnerEnglishApplied AIOpen SourceGAGrace AdeyemiKMKofi Mensah+1SOSam Organiser2.75 (1 of 4)2.7
Designing for Model DeprecationAcceptedApprovedYesAbstractAI-024Every model you build on will be deprecated. Abstraction boundaries that survive a swap, the regression suite that makes a migration a day instead of a quarter, and what not to abstract.Talk (30 min)AdvancedEnglishProductOpen SourceSecurityWCWei ChenSOSam Organiser3.00 (1 of 4)0.7
Scaling Human ReviewAcceptedApprovedYesAbstractAI-022Human review is the bottleneck nobody budgets for. Queue design, reviewer fatigue, inter-rater agreement, and the tooling changes that tripled throughput without changing headcount.Workshop (120 min)BeginnerEnglishInfrastructureAgentsCostIKIdris KhanSOSam OrganiserOct 11, 2026, 6:00 PMRoom B4.25 (1 of 4)2.9
What We Learned Serving 10B TokensAcceptedApprovedYesAbstractAI-023Ten billion tokens through one API. Capacity planning, the failure modes that only appear under sustained load, and the three architectural decisions we would make differently.Talk (30 min)IntermediateEnglishEvaluationCostSecuritySRSofia RossiSOSam OrganiserOct 11, 2026, 6:30 PMRoom B2.50 (1 of 4)2.6
Small Models, Big WinsAcceptedApprovedYesAbstractAI-020A 3B model beats a frontier model on our task, at a fraction of the cost. How we found that out, how we validated it, and the distillation pipeline that keeps it current as the task drifts.Talk (30 min)IntermediateEnglishProductOpen SourceTBTomas BergSOSam OrganiserOct 10, 2026, 5:00 PMRoom A2.75 (1 of 4)1.6
Retrieval Quality Is a Data ProblemAcceptedApprovedYesAbstractAI-021Retrieval failures are almost never the embedding model. Chunking, metadata, document freshness and access control — a walk through the four data problems that caused most of our bad answers.Talk (30 min)AdvancedEnglishApplied AIAgentsRAGNFNaomi FischerWCWei ChenSOSam OrganiserOct 12, 2026, 5:00 PM3.25 (1 of 4)1
Testing Nondeterministic PipelinesAcceptedApprovedYesAbstractAI-018You cannot assert equality on a generative pipeline. Property-based checks, golden sets, statistical gates in CI, and how we stopped flaky tests from training the team to ignore red builds.Talk (30 min)AdvancedEnglishInfrastructureCostSecurityLMLuis MoreauSOSam OrganiserOct 10, 2026, 5:00 PMRoom A2.81 (4)1.4
The Case Against ChatbotsAcceptedApprovedYesAbstractAI-019Chat is a lazy default. Three products where we replaced a conversational interface with something more constrained, the adoption change that followed, and the one case where chat genuinely won.Talk (30 min)BeginnerEnglishEvaluationSecurityPRPriya RamanSOSam OrganiserOct 10, 2026, 5:00 PMRoom A2.88 (4)0.6
From Notebook to Production in a WeekAcceptedApprovedYesAbstractAI-017A team of two took an internal prototype to production in six days. What we skipped deliberately, what that cost us later, and the checklist we now use to decide which corners are safe to cut.Talk (30 min)IntermediateEnglishApplied AIAgentsCostAOAda OkonkwoTBTomas BergSOSam OrganiserOct 11, 2026, 4:30 PMMain Hall3.13 (4)0.6
Latency Budgets End to EndAcceptedApprovedYesAbstractAI-016Latency is a chain, not a number. Tracing a request from browser to model and back, finding the three places we were wasting a second each, and the budget we now hold teams to.Talk (30 min)BeginnerEnglishProductAgentsRAGADAmara DialloSOSam OrganiserOct 12, 2026, 9:00 PMMain Hall2.25 (4)2.2
Multi-tenant Isolation for AI WorkloadsAccept QueueApprovedAbstractAI-014Serving many customers from shared model infrastructure without leaking context between them. Isolation boundaries, noisy-neighbour throttling, and the audit story that survived a security review.Talk (30 min)IntermediateEnglishInfrastructureOpen SourceSecuritySWSam WhitfieldSOSam OrganiserOct 10, 2026, 4:00 PMMain Hall3.25 (4)1.7
Guardrails That Do Not Annoy UsersAcceptedApprovedYesAbstractAI-015Safety filters that block real users are a product bug. We describe the layered approach we use, how we measured false positives against actual traffic, and the categories where we deliberately accept more risk.Workshop (120 min)AdvancedEnglishEvaluationOpen SourceRAGYTYuki TanakaSOSam OrganiserOct 10, 2026, 4:00 PMMain Hall3.13 (4)1.4
Caching Strategies for Token SpendAccept QueueApprovedAbstractAI-013Semantic and exact caching in front of a production assistant. Hit rates by traffic type, the staleness bugs that followed, and the invalidation strategy we ended up with after two rewrites.Talk (30 min)BeginnerEnglishApplied AICostSecurityEPElena PetrovaADAmara DialloSOSam Organiser2.31 (4) · 1.25–3.250.5
When to Say No to an AgentAccept QueueApprovedAbstractAI-012Not every problem deserves an agent. A decision framework drawn from six internal projects, three of which we cancelled, and what the cancelled ones had in common that we did not see at the time.Talk (30 min)AdvancedEnglishProductAgentsKMKofi MensahSOSam Organiser3.13 (4)2.9
Building an Internal Model GatewayAccept QueueApprovedAbstractAI-011We put every model call behind one internal service. This covers routing, per-team quotas, the audit trail that made compliance stop worrying, and the migration that let us swap a vendor in an afternoon.Talk (30 min)IntermediateEnglishEvaluationAgentsRAGMCMira CastellanosSOSam OrganiserMay 14, 2027, 5:00 PMOverflow Room2.88 (4)3
Structured Output Without TearsAccept QueueApprovedAbstractAI-010Getting reliable JSON out of a model that would rather write prose. Grammar-constrained decoding, schema repair, and why we stopped asking the model to apologise for malformed output and started making it impossible.Talk (30 min)BeginnerEnglishInfrastructureOpen SourceJHJonas HalvorsenAOAda OkonkwoSOSam OrganiserMay 14, 2027, 4:00 PMOverflow Room2.44 (4) · 1.75–3.753
Edge Inference on a BudgetPendingApprovedAbstractAI-009Running models close to users without a GPU budget. Quantisation choices, the memory ceiling on commodity edge hardware, and an honest account of the quality we traded away and where it turned out to matter.Talk (30 min)AdvancedEnglishApplied AIOpen SourceSecurityGAGrace AdeyemiKMKofi MensahSOSam Organiser2.13 (4)1.2
Streaming UIs for Slow ModelsPendingApprovedAbstractAI-008When the model takes eleven seconds, the interface is the product. Streaming, skeletons, optimistic rendering and the moment users decide something is broken — with the abandonment numbers that changed our minds.Workshop (120 min)IntermediateEnglishProductCostWCWei ChenSOSam Organiser2.88 (4) · 1.75–4.252.8
Prompt Injection in the WildPendingApprovedAbstractAI-007A field report on prompt injection attempts against a public-facing assistant: what was tried, what worked, and which mitigations were theatre. Includes the payload that got past three layers of filtering.Talk (30 min)BeginnerEnglishEvaluationAgentsCostSRSofia RossiSOSam Organiser3.25 (3)1.5
Observability for Nondeterministic SystemsPendingApprovedAbstractAI-006Traditional observability assumes the same input gives the same output. We cover the tracing schema we settled on, how we sample when every request is unique, and how to alert on quality drift without drowning in false positives.Talk (30 min)AdvancedEnglishInfrastructureRAGIKIdris KhanSOSam Organiser3.31 (4) · 2.25–4.52.6
Fine-tuning Is Not the AnswerPendingApprovedAbstractAI-005Fine-tuning is the first thing people reach for and usually the wrong one. We compare it against retrieval, prompt work and routing on the same three tasks, with the training costs and the maintenance burden included honestly.Talk (30 min)IntermediateEnglishApplied AIOpen SourceRAGNFNaomi FischerWCWei ChenSOSam Organiser4.00 (4)0.6
The Cost Curve of InferencePendingApprovedAbstractAI-004Inference costs do not scale the way finance expects. A breakdown of where our spend actually went across a year — prefill versus decode, cache hit economics, and the surprising fraction consumed by retries and abandoned streams.Talk (30 min)BeginnerEnglishProductOpen SourceSecurityTBTomas BergSOSam Organiser2.63 (4)1.9
Vector Databases at a Billion RowsPendingApprovedAbstractAI-002Everything is fast at ten million rows. We walk through what actually broke between one hundred million and a billion: index build times, memory-mapped segment churn, and the recall cliff nobody warns you about when you quantise too aggressively.Talk (30 min)IntermediateEnglishInfrastructureAgentsLMLuis MoreauSOSam Organiser3.06 (4)1.2
Evaluating RAG: Beyond VibesPendingApprovedAbstractAI-003Most RAG evaluation is a demo and a feeling. We describe the offline harness we built, why we abandoned answer-similarity scoring, and how a small hand-labelled set of two hundred questions caught regressions our automated metrics happily approved.Talk (30 min)AdvancedEnglishEvaluationCostSecurityPRPriya RamanSOSam Organiser3.00 (4) · 1–52.8
Shipping LLM Agents Without Losing SleepPendingApprovedAbstractAI-001We ran agents in production for eighteen months and most of what we believed at the start was wrong. This covers the retry semantics, the budget guards and the three incidents that shaped our current design, including the one that took a weekend to unpick.Workshop (120 min)BeginnerEnglishApplied AIRAGAOAda OkonkwoTBTomas BergSOSam Organiser3.31 (4)1.1
148 of 48
1 / 1
48 abstracts matching the current filters