When I first integrated a large language model into an AI Android app, I made a rookie mistake: I kept feeding the entire conversation history into the model without managing context windows. The result? Sluggish performance, ballooning API costs, and frustrated users staring at loading spinners.

After shipping AudioBook AI with 50K+ users and building multiple LLM-powered features at CodeBrew Labs, I learned that context window management isn't optional—it's the difference between a snappy AI app and one that feels broken. This post covers the real strategies I've used to handle token limits, optimize inference speed, and keep on-device AI models responsive.

Why Context Window Management Matters

A context window is the maximum amount of text an LLM can "see" at once. GPT-4 has 128K tokens. Llama 2 has 4K. Most on-device models running on Android have even tighter budgets—sometimes just 2K tokens.

Here's why this matters for your machine learning mobile app:

  • Memory pressure: Loading a 7B parameter model on a mid-range Android device leaves maybe 2–3GB free. Keeping long conversations in memory drains that fast.
  • Latency: Longer context = slower token generation. A user asking a question and waiting 10 seconds for a response isn't acceptable.
  • Cost: If you're using cloud-based LLM APIs (like OpenAI or Claude), you pay per token. A 100K token context window costs 50x more than a 2K window.
  • Quality degradation: Models perform worse with extremely long contexts. Relevant information gets lost in noise.

📖 Real Numbers

In my AudioBook AI note-taker, I reduced average inference time from 3.2s to 0.8s by implementing smart context windowing. That's a 4x speedup with zero model changes.

Understanding Token Limits & Memory Constraints

Before optimizing, you need to know what you're working with. On Android, your constraints are:

Device Memory

A Snapdragon 8 Gen 2 phone typically has 8–12GB RAM. Your app gets maybe 4–6GB before the system kills it. Load a 7B parameter model (needs ~14GB in float32, ~7GB in int8), and you're already in trouble.

Token Limits

Every LLM has a maximum context length. The problem: you don't get to use all of it. In practice:

  • Reserve 20–30% for the model's response (output tokens).
  • Account for system prompts and instruction overhead (usually 200–500 tokens).
  • That leaves maybe 60–70% for actual conversation history.

On a 2K token model: 2,000 × 0.65 = 1,300 tokens for conversation. That's roughly 5,000 characters or 800 words. Not much.

Quantization Impact

Quantizing your model to int8 or fp16 cuts memory in half but affects token generation speed. I've found int8 is the sweet spot for Android—you get ~10% accuracy loss and 2x memory savings.

⚠️ Don't Assume Full Context

Just because a model supports 4K tokens doesn't mean you should use all 4K on Android. Test actual device memory before deploying. I've seen apps crash in production because they didn't account for system memory pressure.

The Sliding Window Strategy

The most practical approach I've used is the sliding window: keep only the most recent messages in context, dropping older ones as the conversation grows.

How It Works

  • Define a token budget (e.g., 1,200 tokens for conversation).
  • Store the full conversation locally (SQLite or Firebase Firestore).
  • When preparing input for the LLM, start with the system prompt and most recent messages.
  • Add older messages until you hit the token budget.
  • Discard anything beyond the budget.

Why This Works

Recent messages are most relevant. Users rarely expect the AI to remember conversations from 20 messages ago. By keeping the last 5–10 messages, you preserve conversational coherence while staying within memory limits.

I used this in AI NoteTaker: instead of including the entire note history, I kept only the last 8 notes (usually 800–1,000 tokens). Users never noticed, and inference time stayed under 1 second.

Alternative: Importance Scoring

For higher accuracy, you can score messages by relevance. Use embeddings (smaller models like MPNet run on Android) to find semantically similar past messages and include those instead of just recency.

"The sliding window approach feels simple, but it's deceptively powerful. I've shipped it in production apps with tens of thousands of users, and it never causes complaints about context loss."

Practical Implementation in Kotlin

Here's a real implementation pattern I use for on-device AI with context management:

data class Message(
    val id: String,
    val role: String,  // "user" or "assistant"
    val content: String,
    val tokenCount: Int,
    val timestamp: Long
)

class ContextWindowManager(
    private val maxContextTokens: Int = 1200,
    private val reservedOutputTokens: Int = 300,
    private val systemPromptTokens: Int = 150
) {
    
    private val tokenCounter = TokenCounter()  // Use tokenizers lib
    
    suspend fun buildContextWindow(
        conversationHistory: List<Message>,
        systemPrompt: String
    ): String {
        val availableTokens = maxContextTokens - reservedOutputTokens - systemPromptTokens
        var contextBuilder = StringBuilder()
        contextBuilder.append(systemPrompt).append("\n\n")
        
        var tokensUsed = systemPromptTokens
        
        // Start from most recent and work backwards
        for (message in conversationHistory.asReversed()) {
            val messageTokens = message.tokenCount
            
            if (tokensUsed + messageTokens > availableTokens) {
                // Exceeded budget; stop adding messages
                break
            }
            
            val formattedMessage = when (message.role) {
                "user" -> "User: ${message.content}"
                "assistant" -> "Assistant: ${message.content}"
                else -> message.content
            }
            
            contextBuilder.insert(
                systemPrompt.length + 2,
                "$formattedMessage\n\n"
            )
            tokensUsed += messageTokens
        }
        
        return contextBuilder.toString()
    }
    
    fun estimateTokens(text: String): Int {
        // Rough estimate: 1 token ≈ 4 characters
        // For production, use actual tokenizer
        return (text.length / 4) + 1
    }
}

// Usage in your AI ViewModel
class AINoteTakerViewModel(
    private val contextManager: ContextWindowManager,
    private val llmInference: LLMInference
) : ViewModel() {
    
    suspend fun generateAIResponse(userMessage: String) {
        val conversation = fetchConversationHistory()  // From Room/Firestore
        
        val contextWindow = contextManager.buildContextWindow(
            conversationHistory = conversation,
            systemPrompt = "You are a helpful note-taking assistant."
        )
        
        val response = llmInference.generateText(contextWindow)
        saveResponse(response)
    }
}

This pattern ensures:

  • You never exceed the token budget.
  • Recent messages are prioritized.
  • The system prompt and output space are reserved.
  • You can easily swap the token-counting logic for a real tokenizer.

Real-World Patterns from Production

Pattern 1: Tiered Context Strategies

Different features need different context depths. In my ERP app, the invoice AI assistant needed deep context (full order history), while the chat quick-replies needed minimal context (last 2 messages).

I built a tiered system:

  • Tier 1 (Minimal): Last message only. For quick replies. 200 tokens.
  • Tier 2 (Standard): Last 5 messages + system prompt. For most features. 1,000 tokens.
  • Tier 3 (Deep): Last 15 messages + relevant embeddings. For complex analysis. 2,500 tokens (only on high-end devices).

Detection was automatic based on device memory and feature requirements. Never had a crash.

Pattern 2: Streaming + Context Awareness

Streaming responses feels faster to users and lets you reduce context windows. Instead of the user waiting for a full response and then seeing it, they see tokens appearing in real-time.

I used WebSockets (Firebase Realtime or custom Node.js) to stream responses token-by-token. Users feel snappy performance even with smaller context windows.

Pattern 3: Hybrid Local + Cloud

For critical features, I kept a small on-device model for immediate responses, then called a cloud LLM with full context in the background. User sees instant feedback, and accuracy improves within seconds.

📖 Example

In AudioBook AI, when a user asked for a summary, the on-device Llama 2 (2K tokens) gave a quick preview. Then I streamed an OpenAI response with full document context in the background, updating the UI as it arrived. Users thought it was magic.

Pattern 4: Semantic Compression

Instead of truncating old messages, summarize them. Use a smaller model to condense 10 old messages into 2–3 sentence summary (saves 60% tokens, maintains context).

Implementation:

  • After every 5 user messages, trigger a summarization job.
  • Use a lightweight model (DistilBERT, TinyLlama) to create summaries.
  • Replace old messages with their summaries in the context window.

This is heavier on compute but gives better results for long conversations.

Key Takeaways

  • Context windows are your primary constraint on Android. Token limits are tighter than you think—reserve 30% for output, account for system prompts, and budget realistically for conversation history.
  • The sliding window (recent messages first) is the production-proven pattern. I've shipped it in 6+ apps with 4.5+ star ratings. Users don't miss older context; they care about responsiveness.
  • Implement token counting accurately from day one. Rough estimates fail in production. Use a real tokenizer library and test on target devices before launch.
  • Design multi-tier context strategies. Different features need different depths. Quick replies ≠ analysis tasks. Let your app adapt based on device memory and feature type.
  • Combine context management with streaming for best UX. Small context windows feel instant when responses stream token-by-token. Users care about perceived speed, not raw accuracy.