After 30+ turns, your conversational assistant has noticeably slower responses and sometimes produces less coherent outputs. Investigation shows:
- Average conversations reach 50,000 tokens by turn 35.
- Production logs show that 94% of user messages reference only the preceding 3–5 exchanges.
- The 6% of queries that reference earlier context generally ask about information the user could easily re-state.
Your goal is to improve response speed and quality while preserving a good user experience. What is the most effective approach?
Community Discussion