#ZenDesk
Building IntelliDesk AI: How I Architected a Production-Grade Enterprise ITSM Platform with RAG, WebSockets, and Celery
_By Pruthviraj Janwade_ If you have ever worked in an enterprise environment, you know the dread of filing an IT support ticket. You navigate through a labyrinthine portal, fill out a 12-field form with dropdowns you do not understand, and wait 24 to 48 hours just to receive an email asking: _"Have you tried restarting your machine?"_ Tools like **ServiceNow** , **Jira Service Management** , and **Zendesk** are enterprise powerhouses, but they were built in an era of manual triage and static forms. While large language models (LLMs) have taken consumer tech by storm, enterprise IT service management (ITSM) has remained largely trapped in old paradigms. Over the past few months, I set out to bridge this gap by building **IntelliDesk AI** — a production-grade, AI-powered enterprise ITSM platform featuring a conversational RAG assistant, automated ticket lifecycle management, real-time analytics, and role-based access control. In this deep dive, I want to share the architectural decisions, design patterns, engineering challenges, and lessons learned while building this platform from scratch. ## 1. The Core Philosophy: Conversational-First ITSM The primary goal of IntelliDesk AI was simple: **eliminate the friction of IT support for both employees and agents.** Instead of forcing users to fill out static forms, the primary entry point is **IntelliBot** — an intelligent conversational assistant that: 1. **Understands Natural Language** : An employee simply types, _"My Wi-Fi keeps disconnecting every 10 minutes on the 3rd floor."_ 2. **Performs Semantic Document Search (RAG)** : Retrieves relevant company knowledge base guides with exact source attribution. 3. **Resolves Autonomously** : Walks the employee through step-by-step diagnostic and troubleshooting steps. 4. **Auto-Creates & Escalates Tickets**: If self-service fails, IntelliBot extracts the key details (category, urgency, root symptoms), provisions a ticket in the database, and assigns it to the on-duty IT team — without the user filling out a single form field. ## 2. High-Level System Architecture To ensure modularity, scalability, and maintainability, the system is designed around **Clean Architecture** principles: ┌────────────────────────────────────────────────────────┐ │ Browser (React SPA) │ └───────────────────────────┬────────────────────────────┘ │ (HTTP / WSS) ▼ ┌────────────────────────────────────────────────────────┐ │ NGINX (Reverse Proxy & Rate Limiter) │ └──────────────┬──────────────────────────┬──────────────┘ │ (Proxy API) │ (Proxy WebSocket) ▼ ▼ ┌────────────────────────────────────────────────────────┐ │ Flask 3 REST API + Socket.IO (Eventlet) │ │ Controller ──> Service ──> Repository ──> Model │ └───────┬──────────────┬──────────────┬──────────────┬───┘ │ │ │ │ ▼ ▼ ▼ ▼ ┌──────────────┐┌──────────────┐┌──────────────┐┌──────────────┐ │ PostgreSQL ││ Redis Cache ││ ChromaDB ││ Groq API │ │ (Primary DB) ││ & Broker ││(Vector Store)││(Llama 3.3 70B│ └──────────────┘└──────┬───────┘└──────┬───────┘└──────────────┘ │ │ ▼ │ ┌───────────────┐ │ │ Celery Worker │◄──────┘ │ (Async Tasks) │ (Local Sentence Transformers) └───────────────┘ ### Architectural Highlights * **Strict Layer Separation** : * `Controllers`: Pure HTTP routing, input validation (Marshmallow), and status codes. * `Services`: Business logic, domain rules, and workflow orchestration. * `Repositories`: Database queries abstracted via SQLAlchemy 2.0 ORM. * `Models`: Data definitions and relationship mappings. * **Microservices Orchestration** : Fully containerized using Docker and orchestrated with Docker Compose (Postgres, Redis, ChromaDB, Flask, Celery Worker, Celery Beat, Flower, NGINX). ## 3. Designing the Retrieval-Augmented Generation (RAG) Pipeline A major pitfall of many LLM projects is high hallucination rates and exorbitant API costs. Here is how I addressed both: ### Step 1: Document Processing & Chunking When an IT administrator uploads internal documentation (PDFs, DOCX, TXT): 1. Text is extracted cleanly using `PyPDF2` or `python-docx`. 2. A recursive character text splitter splits documents into chunks of **800 characters with an overlap of 150 characters**. This overlap preserves semantic context across chunk boundaries. ### Step 2: Zero-Cost Local Embeddings Instead of calling paid embedding APIs (e.g., OpenAI `text-embedding-ada-002` or `text-embedding-3-small`), I integrated **`sentence-transformers/all-MiniLM-L6-v2`** directly into the container. * It runs locally on CPU with inference times under 15ms per chunk. * Output dimensionality: 384 vectors. * Result: **$0 embedding cost and zero network overhead.** ### Step 3: Vector Indexing & Semantic Search Embeddings are indexed in **ChromaDB**. When a query comes in: 1. The user's prompt is embedded using the same MiniLM model. 2. ChromaDB runs cosine similarity to fetch the **Top-K most relevant chunks**. 3. A confidence score threshold filters out low-relevance matches to prevent hallucinations. ### Step 4: LLM Generation with Strategy Pattern To avoid vendor lock-in, I implemented an AI provider abstraction using the **Strategy Pattern** : class LLMProvider(ABC): @abstractmethod def generate_stream(self, prompt: str, system_message: str): pass class GroqProvider(LLMProvider): def generate_stream(self, prompt: str, system_message: str): # Ultra-fast inference using Groq Llama 3.3 70B ... Groq’s LPU (Language Processing Unit) delivers inference speeds of **~250-300 tokens/sec** , making real-time streaming feel instantaneous. ## 4. Real-Time Streaming: Replacing Polling with WebSockets A common issue with AI chat interfaces is the delay while waiting for the LLM to complete its full response. Initially, simple REST polling or Server-Sent Events (SSE) were considered, but because IntelliDesk already needed bidirectional communication for live dashboard updates, I chose **Flask-SocketIO with an Eventlet worker**. ### The Streaming Protocol 1. Client emits `ai:chat` with user query and session ID. 2. Server validates authentication token and emits `ai:stream:start`. 3. As Groq yields text tokens, server immediately emits `ai:stream:chunk` payloads. 4. Upon completion, server emits `ai:stream:done` along with formatted source citations and confidence metrics. Client (React) Server (Flask + SocketIO) │ │ │ ─── emit('ai:chat', { prompt, sessionId }) ──> │ │ │ ──> Vector Search (ChromaDB) │ │ ──> Stream from Groq (Llama 3.3) │ <── emit('ai:stream:start') ────────────────── │ │ <── emit('ai:stream:chunk', { token: '1.' }) ──│ │ <── emit('ai:stream:chunk', { token: ' Turn' })│ │ <── emit('ai:stream:done', { citations }) ──── │ This reduced perceived latency from **4–6 seconds down to under 200ms**. ## 5. Background Jobs & Asynchronous Workflows (Celery + Redis) Heavy operations should never block an HTTP request. I set up **Celery 5** with **Redis 7** using dedicated priority queues: Queue | Tasks Handled ---|--- `documents` | PDF parsing, chunking, vector embedding generation `ai` | Background ticket intent classification & summarization `email` | SMTP notifications for ticket status and SLA warnings `reports` | Periodic CSAT, SLA metrics, and analytics compilation By pairing Celery with **Celery Beat** and **RedBeat** , scheduled tasks run continuously in the background (e.g., checking for SLA breach thresholds every 60 seconds). For monitoring, **Flower** provides a visual dashboard of task throughput and worker health. ## 6. Frontend Engineering with React 18 & TypeScript The frontend was built to feel like modern software from Linear or Vercel: * **State Strategy** : Clear separation between server cache and client state: * **TanStack Query (React Query)** handles server synchronization, automatic caching, and background invalidation for tickets and analytics. * **Redux Toolkit** manages local state (active AI chat session, dark/light theme, UI modals). * **Socket Lifecycle Management** : Custom React hooks manage WebSocket connection lifecycles, graceful reconnection, and event buffering to ensure messages aren't lost during page navigation. ## 7. Zero-Dollar Infrastructure: Running Production for $0/Month One of my proudest milestones with this project was achieving enterprise-grade capability on a **$0/month infrastructure footprint** : * **Compute API** : Render Free Tier * **Frontend SPA** : Vercel Global Edge Network * **Primary Database** : Neon Serverless PostgreSQL * **Embeddings** : Local CPU execution (`all-MiniLM-L6-v2`) * **LLM Inference** : Groq Free Developer Tier (Llama 3.3 70B) * **Vector DB** : Embedded ChromaDB instance ## 8. What's Next: Enterprise Kubernetes Deployment on AWS EKS With the application architecture fully validated, I am now moving to the next engineering milestone: **deploying IntelliDesk AI onto an enterprise-grade AWS EKS (Elastic Kubernetes Service) cluster.** ### The AWS Deployment Blueprint: 1. **Infrastructure as Code (IaC)** : Provisioning AWS VPC, subnets, IAM Roles for Service Accounts (IRSA), and managed node groups using **Terraform**. 2. **Kubernetes Packaging** : Writing modular **Helm Charts** for the Flask backend, Celery workers, and NGINX Ingress Controller. 3. **Managed Services Integration** : * Amazon RDS (PostgreSQL Multi-AZ) * Amazon ElastiCache (Redis) * Amazon S3 for secure document storage 4. **GitOps Continuous Delivery** : Automated cluster reconciliation and deployments using **ArgoCD**. 5. **Observability** : **Prometheus** for cluster metrics and **Grafana** for executive operational dashboards. ## 9. Conclusion & Takeaways Building IntelliDesk AI taught me several core engineering lessons: * **Design for abstraction early** : Isolating the LLM provider behind a clean interface saved hours when switching between models. * **RAG is only as good as chunking** : Tuning chunk overlap and semantic boundaries matters far more than simply picking a larger LLM. * **User experience is latency-bound** : Real-time streaming via WebSockets fundamentally transforms conversational AI from a sluggish utility into a delightful product. ### Explore the Code The entire project is open-source under the MIT license: * **GitHub Repository** : github.com/Pruthviraj-333/intellidesk-ai * **Watch the Video Walkthrough** : YouTube Demo _If you found this breakdown valuable, feel free to star the repo or connect with me on LinkedIn!_
dev.to
October 1, 2026 at 11:58 AM
Or there are no routes to get feedback at all. Or there technically are routes to submit feedback, but either it requires specific technical knowledge (github), or has never shown a sign of even being looked at (the zendesk or whatever form).
September 30, 2026 at 11:42 PM
#nomanssky patche 7.05 www.nomanssky.com/2026/09/cosm...
Celui là il va faire du bien "Fixed an issue that prevented the gravitino coil from being holstered' :) #cosmos #hellogames
Cosmos 7.05 - No Man's Sky
Hello Everyone, Thank you to everyone playing No Man&#x2019;s Sky &#x2013; Cosmos, especially those taking the time to report any issues they encounter via Zendesk or console crash reporting. We...
www.nomanssky.com
September 30, 2026 at 3:10 PM
I'm just 😭 because how long will this take? I put in through their zendesk thingy but I'm not sure if that's gonna get me a faster response than email
September 29, 2026 at 3:48 PM
Confused by personal AI agents? Here's some inspiration on how to use them from top executives at companies like Anduril and Adobe. https://bit.ly/4hTuaTG
Here are the tasks that execs from Anduril to Zendesk are giving their personal agents
Top execs at companies like Anduril, Adobe, and KPMG are using personal AI agents for tasks like booking dog walkers and flu shots.
www.businessinsider.com
September 29, 2026 at 11:50 AM
→ Nextjs app and Auth via Auth0 for deployment

Connectors working today: GitHub Issues, Airtable, Tally Forms, CSV/notes, and visual files.

Enterprise connectors like (Zendesk, HubSpot, Salesforce, Slack, Intercom, Jira) coming soon

Repo → github.com/Studio1-OSS...
GitHub - Studio1-OSS/revenue-intelligence: Customer revenue intelligence with evidence-backed AI insights. Powered by Nebius Token Factory, Qwen3.8-27B, Turso, and Auth0. Bring your own AI key.
Customer revenue intelligence with evidence-backed AI insights. Powered by Nebius Token Factory, Qwen3.8-27B, Turso, and Auth0. Bring your own AI key. - Studio1-OSS/revenue-intelligence
github.com
September 28, 2026 at 9:32 PM
Don't pay for Zendesk, use Chatwoot

(SAVE THIS before it disappears)

#onlinebusiness #aitools #Productivity
September 28, 2026 at 7:41 PM
en.wikipedia.org/wiki/Zendesk

ZENDESK YOU'LL FIND AT BOTTOM OF WEBSITES
Zendesk - Wikipedia
en.wikipedia.org
September 28, 2026 at 4:38 PM
Zoho Desk's cheapest paid tier includes AI ($7/user/mo); Zendesk holds AI back until Suite Team at $55/agent/mo. Where each one paywalls AI and phone:

https://itsupport.aramagio.com/blog/zoho-desk-vs-zendesk-small-business/?utm_source=social&utm_medium=bluesky&utm_campaign=2026-W40-02
September 28, 2026 at 2:02 PM
Akkio - MonkeyLearn 17. Customer Support - Intercom AI - Zendesk AI - Tidio - Freshdesk AI - Forethought 18. Sales - Apollo AI - Gong - Clay - Lavender - http://Reply.io 19. Finance - http://Vic.ai - Truewind - Zeni - Pilot AI - Grid AI 20. HR & Hiring - HireEZ - Pymetrics -
September 27, 2026 at 6:59 PM
Lol my job literally morphed 6 months ago into assisting in creating and testing a zendesk chatbot and I hate my life. Users are so frustrated and it's not reducing costs in any way because everything still becomes a ticket.
September 27, 2026 at 11:04 AM
This is how you use linkedIn when your industry is no longer employing, right?
lnkd.in/p/eqx4bpUr
#hostelworld #customerservice #zendesk #ai #auckland | Ryan Creedon
Forwarded to #hostelworld #customerservice powered by #zendesk and partially #ai This is the last of a long email chain of drama with 234 backpackers in #auckland that no one seems to want to address....
lnkd.in
September 27, 2026 at 2:18 AM
no more zendesk. society has moved past the need for zendesk
September 25, 2026 at 3:31 PM
AI agents are no longer just tools—they're becoming full-fledged team members. But how do you manage, train, and pay them? New insights from CIOs, UiPath, and Zendesk show that AI needs clear roles, human oversight, and outcome-based compensation. As AI employees evolve, so must our approaches to ma
September 25, 2026 at 1:00 PM
🚀 New remote Customer Support role
Level 2 Support Representative — HappyCo
🌎 USA Only · 💵 $60K - $120K USD

#RemoteWork #RemoteJobs
Level 2 Support Representative
HappyCo · USA Only · $60K - $120K USD · support, product, zendesk
remotearmy.io
September 25, 2026 at 10:29 AM
Rails Girls tickets are now live! 💜

On October 10, we’re bringing the Rails Girls community together at Zendesk Pune for a full day of learning, coding, and community.

🎟️ Grab your ticket: deccanqueenonrails.com/tickets

#RailsGirls #DeccanQueenOnRails
#RubyOnRails #RailsCommunity #Pune
September 24, 2026 at 9:30 AM
Amazon Bedrock Managed Knowledge Base now supports Salesforce and Zendesk as native data source connectors

https://aws.amazon.com/about-aws/whats-new/2026/09/amazon-bedrock-managed-knowledge-base-salesforce-zendesk-native-data-source-connectors/
September 23, 2026 at 10:20 PM
🆕 AWS adds Salesforce and Zendesk connectors to Amazon Bedrock Managed Knowledge Base, enabling direct syncing of articles and posts for enhanced retrieval-augmented generation (RAG) service. For details, see the Amazon Bedrock User Guide.

#AWS #AmazonBedrock #Aiml
Amazon Bedrock Managed Knowledge Base now supports Salesforce and Zendesk as native data source connectors
AWS announces Salesforce and Zendesk data source connectors for Amazon Bedrock Managed Knowledge Base, a fully managed retrieval-augmented generation (RAG) service. Customers can now sync Salesforce knowledge articles and Zendesk articles and community posts directly into their managed knowledge base. Salesforce data source connector and Zendesk data source connector in the Amazon Bedrock User Guide. For more information about Amazon Bedrock Managed Knowledge Base, visit the Amazon Bedrock Knowledge Bases product page.
aws.amazon.com
September 23, 2026 at 8:10 PM
Amazon Bedrock Managed Knowledge Base now supports Salesforce and Zendesk as native data source connectors

AWS announces Salesforce and Zendesk data source connectors for Amazon Bedrock Managed Knowledge Base, a fully managed retrieval-augmented generation (RAG) servic...

#AWS #AmazonBedrock #Aiml
Amazon Bedrock Managed Knowledge Base now supports Salesforce and Zendesk as native data source connectors
AWS announces Salesforce and Zendesk data source connectors for Amazon Bedrock Managed Knowledge Base, a fully managed retrieval-augmented generation (RAG) service. Customers can now sync Salesforce knowledge articles and Zendesk articles and community posts directly into their managed knowledge base. Previously, bringing content from these platforms into Bedrock Knowledge Bases required building custom ingestion pipelines—now, you provide your instance credentials, and the connectors handle data crawling, metadata extraction, and incremental sync automatically. These connectors make it easy to build AI agents and assistants grounded in the support and product knowledge your teams already maintain in Salesforce and Zendesk. For example, power a customer-facing support bot with up-to-date Zendesk help center articles and community answers, or build an internal sales enablement assistant that retrieves relevant Salesforce knowledge articles during deal preparation. By keeping your knowledge base in sync with these platforms, your retrieval-augmented generation applications always reflect the latest content without manual intervention. To learn more, see https://docs.aws.amazon.com/bedrock/latest/userguide/kb-managed-ds-salesforce.html and https://docs.aws.amazon.com/bedrock/latest/userguide/kb-managed-ds-zendesk.htmlin the Amazon Bedrock User Guide. For more information about Amazon Bedrock Managed Knowledge Base, visit the https://aws.amazon.com/bedrock/knowledge-bases/.
aws.amazon.com
September 23, 2026 at 8:05 PM
Amazon Bedrock Managed Knowledge Base now supports Salesforce and Zendesk as native data source connectors
AWS announces Salesforce and Zendesk data source connectors for Amazon Bedrock Managed Knowledge Base, a fully managed retrieval-augmented generation (RAG) service. Customers can now sync Salesforce knowledge articles and Zendesk articles and community posts directly into their managed knowledge base. Previously, bringing content from these platforms into Bedrock Knowledge Bases required building custom ingestion pipelines—now, you provide your instance credentials, and the connectors handle data crawling, metadata extraction, and incremental sync automatically. These connectors make it easy to build AI agents and assistants grounded in the support and product knowledge your teams already maintain in Salesforce and Zendesk. For example, power a customer-facing support bot with up-to-date Zendesk help center articles and community answers, or build an internal sales enablement assistant that retrieves relevant Salesforce knowledge articles during deal preparation. By keeping your knowledge base in sync with these platforms, your retrieval-augmented generation applications always reflect the latest content without manual intervention. To learn more, see Salesforce data source connector and Zendesk data source connector in the Amazon Bedrock User Guide. For more information about Amazon Bedrock Managed Knowledge Base, visit the Amazon Bedrock Knowledge Bases product page.
dlvr.it
September 23, 2026 at 8:04 PM
Amazon Bedrock Managed Knowledge Base now supports Salesforce and Zendesk as native data source connectors

AWS finally lets you plug Salesforce into Bedrock without building your own pipeline. Only took them until 2025 to add a feature Zapier's had for a decade. Pricing details? LOL no.
September 23, 2026 at 8:04 PM