Overview of Inside LinkedIn's cognitive memory agent for agentic personalization
This episode of the Stack Overflow Podcast features Praveen Bhattagutla, Principal AI Researcher at LinkedIn, discussing how LinkedIn built a cognitive memory agent to power more personalized, stateful recruiter experiences. The conversation focuses on how LinkedIn captures, organizes, retrieves, and governs memory across recruiter interactions so AI agents can better understand preferences, reduce friction, and improve hiring workflows at scale.
Why LinkedIn Built a Memory Agent
LinkedIn’s recruiter-facing hiring assistant already helped manage hiring workflows, but the team noticed an important pattern: recruiters repeatedly expressed preferences, refined job criteria, and gave candidate feedback over time. Those signals were sticky and reusable across sessions, roles, and related LinkedIn products.
The memory agent was built to:
- give AI agents a true sense of state
- personalize recruiter experiences more deeply
- manage the full memory lifecycle rather than just retrieve context
- make recruiter workflows more efficient and less repetitive
The Four Layers of Memory
Praveen described LinkedIn’s memory system as a layered structure, each serving a different purpose:
1. Conversation Memory
- Captures the most recent interaction context
- Reflects what the user is doing right now
- Helps the agent stay aligned with the current task
2. Episodic Memory
- Stores recent, time-bound interactions with provenance
- Supports temporal queries like “what happened most recently?”
- Helps trace preferences back to specific activities
3. Semantic Memory
- Aggregates longer-term preferences and knowledge
- Pulls signals across sessions and across related LinkedIn surfaces
- Helps personalize the overall experience based on durable patterns
4. Procedural Memory
- Captures the “how” of a user’s behavior
- Reflects the user’s decision-making style and tradeoffs
- Useful for understanding how different recruiters approach similar hiring tasks
Key Engineering Challenges
Managing a High-Volume Interaction Stream
Recruiter workflows are multi-step and often interleaved:
- writing and refining job descriptions
- reviewing candidates
- giving feedback
- archiving, shortlisting, or reaching out
That creates challenges around:
- identifying session boundaries
- compacting context without losing signal
- preserving discoverability and provenance
- keeping memory fresh as preferences change over time
Handling Conflicting or Changing Preferences
Users may change their minds mid-process, such as:
- shifting locations
- changing preferred skills
- re-prioritizing seniority or workplace type
The system must:
- prioritize recent information appropriately
- resolve conflicts
- keep latency low
- avoid surfacing stale preferences
Hierarchical Memory Structure
For recruiters, LinkedIn found a tree-like hierarchy worked well:
- leaf level: preferences for a specific hiring project
- recruiter level: aggregated preferences across that recruiter’s work
- cohort/company level: shared patterns across recruiters working in similar roles or within the same organization
This structure helps:
- bootstrap new recruiters with company-specific hiring patterns
- share useful domain intelligence within a cohort
- preserve personalization while still enabling collaboration
Platform Design vs. Application Customization
The memory platform is designed to be reusable across applications, not just recruiting.
Standardized Platform Components
- common ingestion and retrieval APIs
- shared orchestration/reasoning layer
- generic tools for memory access
- consistent synthesis and query handling
Customizable Parts
- memory structure can vary by application
- one app may use a tree, another a graph
- episode boundaries may differ depending on the task
- some memories are inferred, others explicitly stored
The core idea: standardize the platform, customize the memory representation where needed.
Relevance, Retrieval, and Transparency
Praveen split relevance into two major problems:
- What should be ingested into memory?
- What should be retrieved for a given query?
LinkedIn’s approach emphasizes:
- domain expertise in deciding what matters
- provenance and typed citations for transparency
- memory records that explain why a preference was inferred
- the ability for users to correct or delete preferences later
This is meant to prevent the system from becoming a black box and to reduce user friction when inferred preferences are wrong.
Privacy, Isolation, and Access Control
Because LinkedIn handles sensitive professional data, the system is built with strong safeguards:
- isolated storage per application
- access tied to authenticated user credentials
- ownership tags on stored memory
- credential propagation through every workflow step
- restricted access to only the data a user is allowed to see
The result is personalized memory with controlled sharing where appropriate, such as within an approved company cohort.
Performance and Cost Optimizations
A major theme in the episode was reducing unnecessary AI work.
What Changed
LinkedIn moved away from slower, more expensive approaches like GraphRAG-style rebuilding because they required too many LLM calls and didn’t scale well.
Key Optimizations
- incremental updates to tree nodes instead of rebuilding everything
- parallel planning instead of sequential planning
- selective LLM use only when reasoning is needed
- structured outputs to limit token generation
- prefix caching / chunk refill optimizations at the serving layer
- staleness policies like aging out memory after 6–12 months when appropriate
These optimizations help fit memory into a small slice of the total latency budget.
Evaluation and Success Metrics
LinkedIn uses a multi-tier evaluation framework to ensure memory quality.
Important goals include:
- preserving the original entities and meaning from input
- maintaining provenance and citations
- minimizing context loss
- reducing friction and repeated user corrections
- improving recruiter helpfulness and product impact
A major downstream signal is whether memory reduces repeated work or causes more correction, which would indicate poor recall or bad prioritization.
Future Directions
Praveen highlighted several active areas of research and development:
- better memory abstractions beyond fixed tool interfaces
- more flexible storage backends, possibly including virtual file systems
- improved attribution and evaluation methods
- dynamic identification of session boundaries
- end-to-end optimization across the full memory lifecycle
- adapting memory systems as LLMs and deep agents become more capable
Notable Takeaways
- Memory for agentic systems is more than conversation history; it needs structure, provenance, freshness, and governance.
- A layered memory model can support both immediate task context and long-term personalization.
- For enterprise-scale systems like LinkedIn, latency, cost, and correctness are all first-class constraints.
- Transparency matters: users should be able to understand, correct, and override inferred preferences.
- The future of agentic personalization likely depends on more sophisticated memory infrastructure, not just better prompting.
Additional Notes
- Praveen mentioned that one of the related papers was accepted at KDD.
- LinkedIn plans to share more blog posts and publications as the work continues.
