# Mukesh Bishnoi > Frontend Engineer specialising in React.js, Next.js, TypeScript, Node.js and AWS. Founding Engineer at Foosh AI, building a scalable AI workflow platform. Previously Software Engineer at Unolo and Step Security, and open source developer at Coasts.dev (YC '25), Nightwatch.js and Empirical. Founding Engineer based in Pune, India. Website: https://www.mukeshbishnoi.com Founding Engineer building AI-powered products end-to-end. From React dashboards to LLM agents, I ship fast and break nothing. ## Contact and Profiles - Email: mukeshb.work@gmail.com - GitHub: https://github.com/mukeshblackhat - LinkedIn: https://www.linkedin.com/in/mukesh-bishnoi/ - X: https://x.com/01Mukesh29 - Résumé (PDF): https://www.mukeshbishnoi.com/resume.pdf These profiles all belong to Mukesh Bishnoi. The GitHub account mukeshblackhat is the source of the open source work and personal projects listed below; the public contribution history is on that profile. ## Work Experience ### Foosh AI — Founding Engineer April 2025 – Present · Remote · https://foosh.ai Building end-to-end AI-powered products as the founding engineer. Architecting full-stack solutions using Next.js and TypeScript, integrating LLM agents and AI pipelines to solve complex business problems from ideation to production. - Architected a scalable AI workflow platform using React Flow, TypeScript, React Query, and AWS Step Functions with Lambda orchestration, enabling users to visually compose and execute workflows across 10+ AI models for image and video generation while supporting robust state management and asynchronous API execution. - Engineered a high-performance media library handling thousands of images and videos by implementing CDN delivery, multi-layer caching, lazy loading, virtualization, and paginated fetching, reducing initial load time by 60%+ and significantly improving rendering performance for large datasets. - Optimized backend performance by restructuring AWS resource usage (Lambda, Step Functions, S3) and streamlining API orchestration across multiple third-party AI platforms, reducing execution latency and improving reliability of asynchronous workflow processing at scale. Technologies: React Flow, TypeScript, React Query, AWS Step Functions, AWS Lambda, Amazon S3, Next.js, CDN ### Unolo — Software Engineer Nov 2023 – April 2025 · Remote · https://unolo.com Built internal dashboards to track company performance and sales metrics. Optimized frontend performance through virtualization techniques and led a major initiative to create customizable client-facing dashboards, improving user experience across the platform. - Built an internal analytics dashboard surfacing at-risk customer signals, driving a 50% reduction in churn and significantly strengthening annual revenue retention. - Optimized website rendering and data handling to improve performance by 40% while scaling map capacity from 4,000 to 40,000+ points, boosting user engagement by 25%. - Led frontend development using Next.js, TypeScript, and MUI to deliver key product features, driving a 30% increase in adoption among enterprise clients including Flipkart, Adani, and Tata. Technologies: Next.js, TypeScript, Material-UI, React ### Step Security — Software Engineer August 2023 – Nov 2023 · Remote · https://www.stepsecurity.io/ Led migration of the product from React to Next.js, significantly improving performance and SEO. Set up the deployment pipeline on AWS and optimized the application for faster load times and better search engine visibility. - Migrated codebase from React.js to Next.js, reducing page load times by 40% and increasing SEO rankings by 25%. - Optimized AWS deployment, cutting infrastructure costs by 30% while maintaining 99.9% uptime. - Contributed to product growth, reaching over 3,500 open-source repositories and securing partnerships with Google, Microsoft, Node.js, and Datadog. Technologies: React.js, Next.js, AWS ## Open Source ### coasts.dev — Coasts (YC'25) March 2026 · https://coasts.dev Refactored core modules of coasts.dev (YC'25), reducing code complexity by breaking monolithic logic into small, reusable, and independently testable functions, improving codebase maintainability and test coverage. - Refactored core modules of coasts.dev (YC'25), reducing code complexity by breaking monolithic logic into small, reusable, and independently testable functions. - Improved codebase maintainability and test coverage. Impact: Improved codebase maintainability and test coverage ### Nightwatch.js — BrowserStack April 2024 · https://nightwatchjs.org Contributed features to Nightwatch.js, an open-source test automation framework. - Contributed features to Nightwatch.js, an open-source test automation framework. ### Empirical-Run CLI — Empirical May 2023 · https://www.empirical.run/ Enhanced the Empirical-Run CLI, streamlining configuration and setup for AI prompt testing workflows and reducing setup time by 30%. - Enhanced the Empirical-Run CLI, streamlining configuration and setup for AI prompt testing workflows. - Reduced setup time by 30%. Impact: 30% reduction in setup time ### Stakwork & Sphinx — Stakwork Dec 2023 · https://stakwork.com Resolved 3 critical bugs in the LeetCode Helper Chrome Extension, improving stability for 20,000+ users and increasing user satisfaction by 25%. - Resolved 3 critical bugs in the LeetCode Helper Chrome Extension. - Improved stability for 20,000+ users and increased user satisfaction by 25%. - Feature enhancement and bug fixing. Impact: Stability for 20,000+ users, 25% higher satisfaction ## GitHub Activity Public contribution history: https://github.com/mukeshblackhat The site renders a contribution calendar for mukeshblackhat on the home page. It is drawn client-side from GitHub's own data, so the counts are not in this file — read them from the GitHub profile itself, which is the authoritative source. ## Skills **Frontend:** React.js, Next.js, TypeScript, JavaScript, React Flow, React Query, Redux Toolkit, Zustand, Redux, Material UI, Tailwind CSS, Styled-Components, SCSS, SASS, Bootstrap, HTML5, CSS3 **Backend:** Node.js, Express.js, Python, FastAPI, MongoDB, RESTful APIs, C++ **DevOps & Tools:** AWS, AWS Lambda, AWS Step Functions, AWS DynamoDB, Vercel, Railway, Heroku, Git, GitHub, Performance Optimization, Agile Methodologies, Figma, VS Code, Chrome Web Store ## Projects ### Video Editor AI (2025) AI-powered video background replacement pipeline using Meta's SAM 2 model. Automatically segments subjects and replaces backgrounds in real-time video streams. Technologies: Python, SAM 2, OpenCV, ffmpeg, MatAnyone Source: https://github.com/mukeshblackhat/video-editor ### Browser Bridge MCP (2025) MCP server that bridges Claude Code with browser DevTools for real-time debugging. Enables AI-assisted development by connecting Claude directly to your browser. Technologies: JavaScript, Node.js, MCP, WebSocket, Chrome DevTools Source: https://github.com/mukeshblackhat/browser-bridge-mcp ### PR Analyser (2024) Automated pull request analysis tool that reviews code changes, identifies potential issues, and provides actionable feedback for better code quality. Technologies: JavaScript, Node.js, GitHub API ### Extension Creator CLI (2024) CLI tool to scaffold and generate browser extensions quickly. Streamlines the development workflow for building Chrome and Firefox extensions. Technologies: JavaScript, Node.js, CLI ## Education ### Army Institute of Technology (AIT) Pune, India · Oct 2020 – Jul 2024 ## Achievements - **PMSS Scholarship Holder** (2020-2024) — PMSS Scholarship Holder (2020-2024), awarded to top 1% of engineering students. - **Technical Head, Open Source Software Club** — Technical Head at Open Source Software Club, AIT Pune, leading 20+ successful projects. - **Technical Writing** — Published 15+ technical articles on Medium, garnering 50,000+ views. - **Information Security and Digital Forensics (ISDF)** — Active member of Information Security and Digital Forensics (ISDF), contributing to 3 major security projects. - **Freelance Web Development** — Completed 10+ freelance web development projects with 100% client satisfaction rate. ## FAQ **Who is Mukesh Bishnoi?** Mukesh Bishnoi is a Founding Engineer based in Pune, India. He is a Frontend Engineer specialising in React.js, Next.js, TypeScript, Node.js and AWS. Founding Engineer at Foosh AI, building a scalable AI workflow platform. Previously Software Engineer at Unolo and Step Security, and open source developer at Coasts.dev (YC '25), Nightwatch.js and Empirical. **What does Mukesh Bishnoi do now?** He is Founding Engineer at Foosh AI (April 2025 – Present). Building end-to-end AI-powered products as the founding engineer. Architecting full-stack solutions using Next.js and TypeScript, integrating LLM agents and AI pipelines to solve complex business problems from ideation to production. **Where has Mukesh Bishnoi worked?** Founding Engineer at Foosh AI (April 2025 – Present), Software Engineer at Unolo (Nov 2023 – April 2025) and Software Engineer at Step Security (August 2023 – Nov 2023). **What technologies does Mukesh Bishnoi specialise in?** Frontend — React.js, Next.js, TypeScript, JavaScript, React Flow, React Query, Redux Toolkit, Zustand, Redux, Material UI, Tailwind CSS, Styled-Components, SCSS, SASS, Bootstrap, HTML5, CSS3. Backend — Node.js, Express.js, Python, FastAPI, MongoDB, RESTful APIs, C++. DevOps & Tools — AWS, AWS Lambda, AWS Step Functions, AWS DynamoDB, Vercel, Railway, Heroku, Git, GitHub, Performance Optimization, Agile Methodologies, Figma, VS Code, Chrome Web Store. **What open source has Mukesh Bishnoi contributed to?** coasts.dev (Coasts (YC'25)) in March 2026, Nightwatch.js (BrowserStack) in April 2024, Empirical-Run CLI (Empirical) in May 2023 and Stakwork & Sphinx (Stakwork) in Dec 2023. Refactored core modules of coasts.dev (YC'25), reducing code complexity by breaking monolithic logic into small, reusable, and independently testable functions, improving codebase maintainability and test coverage. Contributed features to Nightwatch.js, an open-source test automation framework. Enhanced the Empirical-Run CLI, streamlining configuration and setup for AI prompt testing workflows and reducing setup time by 30%. Resolved 3 critical bugs in the LeetCode Helper Chrome Extension, improving stability for 20,000+ users and increasing user satisfaction by 25%. **What projects has Mukesh Bishnoi built?** Video Editor AI (2025), Browser Bridge MCP (2025), PR Analyser (2024) and Extension Creator CLI (2024). Video Editor AI: AI-powered video background replacement pipeline using Meta's SAM 2 model. Automatically segments subjects and replaces backgrounds in real-time video streams. Browser Bridge MCP: MCP server that bridges Claude Code with browser DevTools for real-time debugging. Enables AI-assisted development by connecting Claude directly to your browser. PR Analyser: Automated pull request analysis tool that reviews code changes, identifies potential issues, and provides actionable feedback for better code quality. Extension Creator CLI: CLI tool to scaffold and generate browser extensions quickly. Streamlines the development workflow for building Chrome and Firefox extensions. **Where did Mukesh Bishnoi study?** Army Institute of Technology (AIT), Pune, India, Oct 2020 – Jul 2024. **How can I contact Mukesh Bishnoi?** By email at mukeshb.work@gmail.com. He is also on GitHub at https://github.com/mukeshblackhat, LinkedIn at https://www.linkedin.com/in/mukesh-bishnoi/ and X at https://x.com/01Mukesh29. **Where is Mukesh Bishnoi based?** Pune, India. His recent roles have been remote. ## Writing - [Why We Use Zustand with React Flow to Manage State](https://www.mukeshbishnoi.com/blog/context-to-zustand-migration) — Our AI workflow canvas got laggy at 30 nodes. Splitting the React Context did not fix it, and memo did not either. This is why we moved to Zustand, why Redux lost, and why React Flow already made the decision for us. (2026-08-30) - Why use Zustand with React Flow instead of React Context? → https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-1-why-use-zustand-with-react-flow-instead-of-react-context - Why does a React Flow canvas get slow when you add more nodes? → https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-2-why-does-a-react-flow-canvas-get-slow-when-you-add-more-nodes - How do you update one node's data in React Flow without re-rendering every node? → https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-3-how-do-you-update-one-nodes-data-in-react-flow-without-re-rendering-every-node - Is React Context bad for state that changes often? → https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-4-is-react-context-bad-for-state-that-changes-often - Does splitting React Context into smaller providers fix re-render problems? → https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-5-does-splitting-react-context-into-smaller-providers-fix-re-render-problems - Zustand or Redux for a node editor or canvas app? → https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-6-zustand-or-redux-for-a-node-editor-or-canvas-app - Should editing state and execution state be in the same store? → https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-7-should-editing-state-and-execution-state-be-in-the-same-store - What state does React Flow keep internally, and can you read it? → https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-8-what-state-does-react-flow-keep-internally-and-can-you-read-it - How many nodes can React Flow handle before performance becomes a problem? → https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-9-how-many-nodes-can-react-flow-handle-before-performance-becomes-a-problem - [Foosh: From One Prompt Box to Ten Thousand Rows](https://www.mukeshbishnoi.com/blog/foosh-engineering-experience) — Building an AI workflow platform end to end. Why React Context could not keep up with React Flow, how run state moved from memory to DynamoDB to S3, why we started at five parallel rows, and what happened when an LLM started doing the QA. (2026-08-27) - How do you run a node-graph workflow on AWS? → https://www.mukeshbishnoi.com/blog/foosh-engineering-experience#faq-1-how-do-you-run-a-node-graph-workflow-on-aws - Why does a DynamoDB write fail on a large workflow run, and how do you fix it? → https://www.mukeshbishnoi.com/blog/foosh-engineering-experience#faq-2-why-does-a-dynamodb-write-fail-on-a-large-workflow-run-and-how-do-you-fix-it - Why would two parallel nodes lose each other's output in a workflow run? → https://www.mukeshbishnoi.com/blog/foosh-engineering-experience#faq-3-why-would-two-parallel-nodes-lose-each-others-output-in-a-workflow-run - Should you use polling or websockets for long-running AI generations? → https://www.mukeshbishnoi.com/blog/foosh-engineering-experience#faq-4-should-you-use-polling-or-websockets-for-long-running-ai-generations - How do you make an image gallery with thousands of AI-generated images fast? → https://www.mukeshbishnoi.com/blog/foosh-engineering-experience#faq-5-how-do-you-make-an-image-gallery-with-thousands-of-ai-generated-images-fast - [A Year and a Half at Unolo: Four Problems Worth Writing Down](https://www.mukeshbishnoi.com/blog/unolo-engineering-experience) — Churn caught a month early, a landing page cut from 2.5s to 0.5s, a dashboard the customer assembles themselves, and a map that went from crashing at 4,000 points to smooth at 100,000. (2026-08-24) - How do you render 100,000 points on a map without the browser freezing? → https://www.mukeshbishnoi.com/blog/unolo-engineering-experience#faq-1-how-do-you-render-100000-points-on-a-map-without-the-browser-freezing - Why is a map still janky after virtualizing the list of points? → https://www.mukeshbishnoi.com/blog/unolo-engineering-experience#faq-2-why-is-a-map-still-janky-after-virtualizing-the-list-of-points - How do you cut a landing page load time from 2.5 seconds to under a second? → https://www.mukeshbishnoi.com/blog/unolo-engineering-experience#faq-3-how-do-you-cut-a-landing-page-load-time-from-25-seconds-to-under-a-second ## Questions These Posts Answer **Why use Zustand with React Flow instead of React Context?** Because React Flow already stores its own graph state in Zustand. Node positions, viewport, drag state, selection and connections all live in a Zustand store inside the library. If you keep your node data in React Context as well, two systems own the same data in two different shapes and you have to write a transform and a sync layer between them. Putting your state in Zustand removes that layer: React Flow's change handler becomes an action on your own store, both shapes are updated in the same call, and components subscribe per node instead of per provider. Source: [Why We Use Zustand with React Flow to Manage State](https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-1-why-use-zustand-with-react-flow-instead-of-react-context) **Why does a React Flow canvas get slow when you add more nodes?** Usually it is not React Flow, it is how the node data is subscribed to. If every custom node reads from one React Context, or from the whole nodes array, then changing one node re-renders all of them plus the sidebar and any panel reading the same value. On our AI workflow builder that showed up at around 30 to 40 nodes as sticky dragging and stuttering zoom, and profiling showed 20 to 50 re-renders for a single node being moved. The fix is granular subscriptions, so a component only re-renders when the specific slice it reads actually changes. Source: [Why We Use Zustand with React Flow to Manage State](https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-2-why-does-a-react-flow-canvas-get-slow-when-you-add-more-nodes) **How do you update one node's data in React Flow without re-rendering every node?** Keep the per-node data in a Zustand store keyed by node id, and have each custom node component select only its own key. The default approach — call setNodes on the whole array to change one node's data — makes React Flow reconcile the entire array, and any node component reading that array re-renders with it. With a store, updating node 12's config or execution status touches node 12's component only; nodes 1 to 11 never subscribed to it. Source: [Why We Use Zustand with React Flow to Manage State](https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-3-how-do-you-update-one-nodes-data-in-react-flow-without-re-rendering-every-node) **Is React Context bad for state that changes often?** It is the wrong tool for it. React Context has no concept of a partial subscription: any component that consumes a context re-renders whenever any value in that context changes, even if it only reads one field of it. That is fine for state that rarely changes — theme, auth, locale, feature flags — and it falls apart for state written many times a second, like a canvas being dragged. For that you need a store with selector-based subscriptions, such as Zustand, Redux with React-Redux, or Jotai. Source: [Why We Use Zustand with React Flow to Manage State](https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-4-is-react-context-bad-for-state-that-changes-often) **Does splitting React Context into smaller providers fix re-render problems?** Not on its own. We tried it before switching libraries — one provider per node, with memoised consumers — and changing one node still re-rendered the other nodes and their internals. Slicing providers thinner does not change the rule that everything reading a provider re-renders when that provider's value changes, and React.memo cannot help when the value a component is subscribed to is a new object on every update. It is worth trying because it is cheap, but do not plan around it working. Source: [Why We Use Zustand with React Flow to Manage State](https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-5-does-splitting-react-context-into-smaller-providers-fix-re-render-problems) **Zustand or Redux for a node editor or canvas app?** Both can solve the re-render problem, because both support selector-based subscriptions. We chose Zustand for three reasons: creating a store is a few lines with no provider or reducer wiring, selector subscriptions are the default rather than something you have to be disciplined about, and React Flow's own internal store is Zustand, so there is one state model in the app instead of two. Redux earns its boilerplate on larger apps with many teams and strict action logging; for fixing a canvas re-render storm it was a slower, heavier route to the same place. Source: [Why We Use Zustand with React Flow to Manage State](https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-6-zustand-or-redux-for-a-node-editor-or-canvas-app) **Should editing state and execution state be in the same store?** Keep them separate if the user can edit while something is running. We use one Zustand store for editing — nodes, edges, positions, configs — and a second for execution — status, progress and output per node, kept fresh by polling. The reason is a product requirement: you can edit, and even delete, a node while it is still executing. With separate stores the editor drops a deleted node immediately and the execution store just ignores results for nodes that no longer exist. In one store, every polled update would be writing to the object the editor is writing to. Source: [Why We Use Zustand with React Flow to Manage State](https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-7-should-editing-state-and-execution-state-be-in-the-same-store) **What state does React Flow keep internally, and can you read it?** React Flow keeps node positions, the viewport and zoom, drag state, selection and in-progress connections in an internal Zustand store. You can read it directly with its useStore hook for reactive access, or useStoreApi when you need the current value without subscribing. That is one of the practical advantages of using Zustand for your own state too: the library's state and your state are the same kind of thing, rather than one being a black box behind a Context you cannot select into. Source: [Why We Use Zustand with React Flow to Manage State](https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-8-what-state-does-react-flow-keep-internally-and-can-you-read-it) **How many nodes can React Flow handle before performance becomes a problem?** There is no fixed number, because the ceiling is set by how your node components subscribe to data rather than by the library. With every node reading one React Context we felt it at 30 to 40 nodes, especially with an execution running at the same time. After moving the same graph to Zustand with per-node selectors, that lag disappeared and the 300ms position debounce we had added to hide it was deleted. Fix subscriptions first before reaching for virtualisation or node culling. Source: [Why We Use Zustand with React Flow to Manage State](https://www.mukeshbishnoi.com/blog/context-to-zustand-migration#faq-9-how-many-nodes-can-react-flow-handle-before-performance-becomes-a-problem) **How do you run a node-graph workflow on AWS?** We use one generic AWS Step Functions state machine with Lambda doing the work, rather than generating a state machine per workflow. The machine reads the graph, works out which nodes can run now, prepares each node's inputs, invokes the right Lambda, writes the output back and repeats. The graph is data; the machine that walks it never changes, so saving a new canvas does not deploy anything. Source: [Foosh: From One Prompt Box to Ten Thousand Rows](https://www.mukeshbishnoi.com/blog/foosh-engineering-experience#faq-1-how-do-you-run-a-node-graph-workflow-on-aws) **Why does a DynamoDB write fail on a large workflow run, and how do you fix it?** A DynamoDB item cannot exceed 400KB, and a run record holding every node's output passes that more easily than you would expect. The failure is abrupt — the run does not slow down, it refuses to write. The fix is to put oversized node data in S3 and keep a small pointer in the item, then make every path that reads node data resolve that pointer first, so nothing downstream needs to know where the data actually lives. Source: [Foosh: From One Prompt Box to Ten Thousand Rows](https://www.mukeshbishnoi.com/blog/foosh-engineering-experience#faq-2-why-does-a-dynamodb-write-fail-on-a-large-workflow-run-and-how-do-you-fix-it) **Why would two parallel nodes lose each other's output in a workflow run?** Because both write to the same execution record and one silently overwrites the other. Nothing errors, and the run reports success with an output missing. The fix is ordinary optimistic locking: keep a version on the record and write on the condition that the version is still what you read, re-reading and retrying if someone got there first. Source: [Foosh: From One Prompt Box to Ten Thousand Rows](https://www.mukeshbishnoi.com/blog/foosh-engineering-experience#faq-3-why-would-two-parallel-nodes-lose-each-others-output-in-a-workflow-run) **Should you use polling or websockets for long-running AI generations?** We have polled from the beginning and would again on this architecture. On Lambda, holding a socket open means paying for a function that sits doing nothing, and keeping that connection alive across a deliberately stateless system is work. Generations take thirty seconds to five minutes, so nobody can tell the difference between a socket and a poll — the model is what is slow, and a socket would not make it faster. Source: [Foosh: From One Prompt Box to Ten Thousand Rows](https://www.mukeshbishnoi.com/blog/foosh-engineering-experience#faq-4-should-you-use-polling-or-websockets-for-long-running-ai-generations) **How do you make an image gallery with thousands of AI-generated images fast?** Two ordinary fixes did it. Page the data — we fetch 20 at a time and load the next page about 1000 pixels from the bottom, so the whole workspace is never in the page. And stop sending full-size images to a grid: a CDN in front of the bucket returns a 300px thumbnail at around 50KB instead of a 4MB original the browser then shrinks. One warning — a CDN compresses whether you ask it to or not, so downloads have to bypass it or users get a compressed file they think is the original. Source: [Foosh: From One Prompt Box to Ten Thousand Rows](https://www.mukeshbishnoi.com/blog/foosh-engineering-experience#faq-5-how-do-you-make-an-image-gallery-with-thousands-of-ai-generated-images-fast) **How do you render 100,000 points on a map without the browser freezing?** Never hand the map 100,000 markers. We took a map that crashed past roughly 4,000 points to a smooth 100,000 by building a zoom-based cluster hierarchy: fully zoomed out the map is handed a single country roll-up, then state-level clusters, then district-level, and only when you are close in does it get individual points. At every zoom level the map draws a small number of things, and the points that are not legible at that zoom are never sent. Clusters, individual points and sites also live in separate layers, which is what makes the per-zoom rules expressible at all. Source: [A Year and a Half at Unolo: Four Problems Worth Writing Down](https://www.mukeshbishnoi.com/blog/unolo-engineering-experience#faq-1-how-do-you-render-100000-points-on-a-map-without-the-browser-freezing) **Why is a map still janky after virtualizing the list of points?** Because there were two bottlenecks stacked on each other, and virtualizing only removed one. Windowing the React components stopped the page crashing and made it faster, but zooming still dragged. The test that located the real one was to stop passing points to the map at all — the drag vanished, which proved the remaining cost was the number of markers handed to the map itself, not React. Fixing one of two stacked bottlenecks feels like progress and still feels broken; that is the signal to keep digging rather than conclude the library is just slow. Source: [A Year and a Half at Unolo: Four Problems Worth Writing Down](https://www.mukeshbishnoi.com/blog/unolo-engineering-experience#faq-2-why-is-a-map-still-janky-after-virtualizing-the-list-of-points) **How do you cut a landing page load time from 2.5 seconds to under a second?** Four things, and they compound rather than add. Get image sizes down and stop serving the wrong-sized image per breakpoint, which is what was causing the layout shift and the slow paint. Lazy-import the components that are not on screen at first load instead of bundling them into the initial payload. Put assets on a CDN, which matters as soon as your clients are in different regions from your origin. And tree-shake the unused CSS and JS, because a smaller file served from a nearer edge wins twice. Image and bundle work was the LCP and FCP work — they were not separate tasks. Source: [A Year and a Half at Unolo: Four Problems Worth Writing Down](https://www.mukeshbishnoi.com/blog/unolo-engineering-experience#faq-3-how-do-you-cut-a-landing-page-load-time-from-25-seconds-to-under-a-second) ## Pages - [Home](https://www.mukeshbishnoi.com) — profile, experience, open source, skills, projects - [Résumé](https://www.mukeshbishnoi.com/resume) — the same history in résumé form - [FAQ](https://www.mukeshbishnoi.com/faq) — common questions, answered - [Blog](https://www.mukeshbishnoi.com/blog) — long-form engineering write-ups - [RSS](https://www.mukeshbishnoi.com/blog/rss.xml) - [llms-full.txt](https://www.mukeshbishnoi.com/llms-full.txt) — this file plus every post in full --- # Full Articles ## Why We Use Zustand with React Flow to Manage State Published: 2026-08-30 · Category: tech · URL: https://www.mukeshbishnoi.com/blog/context-to-zustand-migration Our AI workflow canvas got laggy at 30 nodes. Splitting the React Context did not fix it, and memo did not either. This is why we moved to Zustand, why Redux lost, and why React Flow already made the decision for us. We're building a workflow builder. A canvas where you drag in a node, point it at a frontier model, feed it some input, and say what shape you want the output in. Nodes connect to nodes. One node's output is the next one's input. Simple on paper. In practice it means a lot of state changing at once, read by a lot of components, sixty times a second while somebody is dragging. Getting that wrong is what made the canvas feel slow. This is the write-up of how we got it right, and why the answer was decided by a library we had already installed.
30–40 nodes before it got laggy
20–50 re-renders to drag one node
300ms debounce we could delete
2 stores, kept separate on purpose
## Two stores, one canvas Before any of the library choice mattered, we had already split the state in two. - **The editing store** — everything about *building* the workflow. Canvas zoom, node positions, node configs, edges, which node is selected. - **The execution store** — everything about *running* it. Which node is executing, what is pending, what finished, and the output of each node. The reason for the split is a product requirement: you can **edit a workflow while it is executing**. If both lived in one blob, every polled execution update would land on the same object the editor is writing to, and the two would fight. ## Round one: React Context We reached for what was closest. Two providers, one per store, everything inside. ```tsx // The shape that seemed reasonable at the time. type EditingState = { nodes: Node[]; edges: Edge[]; viewport: Viewport; selectedNodeId: string | null; nodeConfigs: Record; updateNodeConfig: (id: string, patch: Partial) => void; }; const EditingContext = createContext(null); export function EditingProvider({ children }: { children: ReactNode }) { const [nodes, setNodes] = useState([]); const [edges, setEdges] = useState([]); const [viewport, setViewport] = useState(DEFAULT_VIEWPORT); const [selectedNodeId, setSelectedNodeId] = useState(null); const [nodeConfigs, setNodeConfigs] = useState>({}); const updateNodeConfig = useCallback((id: string, patch: Partial) => { setNodeConfigs((prev) => ({ ...prev, [id]: { ...prev[id], ...patch } })); }, []); // Every consumer re-renders when any one of these changes. const value = useMemo( () => ({ nodes, edges, viewport, selectedNodeId, nodeConfigs, updateNodeConfig }), [nodes, edges, viewport, selectedNodeId, nodeConfigs, updateNodeConfig], ); return {children}; } ``` A node component then did `useContext(EditingContext)` and pulled out the one field it cared about. That last sentence is the whole bug. It pulls one field out **after** it has already subscribed to all of them. This worked fine for a demo-sized workflow. ## The cracks showed at about 30 nodes Once real workflows hit 30–40 nodes, especially with an execution running alongside, you could feel it: - dragging a node felt sticky, like the canvas had resistance - zooming stuttered - the whole thing felt heavy We profiled, and it was not subtle. Changing **one** value on **one** node was re-rendering every other node, the sidebar that holds all the node configs, and every panel and button reading from that context. Measured: 20 to 50 re-renders for a single node moving.
REACT CONTEXT ZUSTAND SELECTORS n1n2n3 n4n5n6 sidebar · inspector · toolbar n1n2n3 n4n5n6 sidebar · inspector · toolbar edit n5 → 7 components re-render edit n5 → 1 component re-renders
Same edit, same canvas. The only difference is how components subscribe. Filled = re-rendered.
This is the classic Context trap: **a component that consumes a context re-renders whenever any value in that context changes**, even if it only reads one field. With 30+ nodes all subscribed to the same provider, one small update rippled out to everything. ## We tried to fix Context before replacing it Worth saying, because "we swapped libraries" is a bad story if you skipped the cheap fix. Two attempts: 1. **Split the single Context into smaller providers** — down to one provider per node, so a node's update should only touch that node's provider. 2. **Wrap consumers in `memo`** so an unchanged subtree bails out. Neither held. Changing one node still re-rendered the other nodes and their internals. That is the nature of Context: anything reading a provider re-renders when that provider's value changes, and `memo` cannot help you when the value the component *is subscribed to* is a new object on every update. You can slice providers as thin as you like, but the moment they share a parent that re-renders, you are back where you started. So we stopped patching and went looking for a store built for this. ## Redux or Zustand We knew what we needed: **selector-based subscriptions**, where a component re-renders only if the specific slice it reads actually changed. Two candidates. | Criteria | Redux (+ React-Redux) | Zustand | | --- | --- | --- | | Setup effort | Store, reducers, actions, provider | One hook, basically done | | Selective re-rendering | Possible, with careful selector discipline | Built in, selector-first by default | | Learning curve for the team | More concepts to onboard onto | Feels like using a hook | | Time to implement, at our stage | High | Low | | Fit for current product stage | Overkill for the problem we had | Matches the complexity we have | Redux could absolutely have solved this. With React-Redux and disciplined selectors, it is the same fix. But we were not looking for an architecture, we were looking to stop a re-render storm, and the machinery was not worth it yet. Zustand gave us most of the benefit for a fraction of the setup. ## The reason that actually decided it Our canvas is **React Flow**. And React Flow's own internal state — node positions, viewport, drag state, selection, connections — **is Zustand**. Not a coincidence we noticed later; it is how React Flow tells you to manage custom node state once you are past a handful of nodes. That matters because of what we were doing wrong, which was not "we picked the wrong library". It was that **two things owned the same data**. React Flow keeps the graph in its store because it draws the canvas. We kept the graph in Context too, in our own shape. So something had to transform between the two, and then keep them agreeing:
WHILE DRAGGING, THIS RUNS 60×/SECOND React Flow store onNodesChange context setState transform + hash setNodes(nextNodesArray) React Flow sees a new array, reports a change, and we go round again.
The loop. The guard flag and the hash check existed only to break it — about 150 lines whose entire job was keeping two copies of the graph agreeing.
To stop the loop you add a guard flag ("this update came from me, ignore it") and a hash check (is this incoming array actually different, or just a new array with the same contents). Which means serialising the whole nodes array on every render, twice. Which is why position updates got debounced by 300ms. Which is why, while you dragged, React Flow painted the node under your cursor immediately and everything reading our state was a third of a second behind. ## What the fix looked like We did not add a state library. We **stopped being the second owner**.

Before

Context → transform → sync → React Flow, and back again. Two shapes, one guard flag, one hash, one debounce.

After

One Zustand store that React Flow talks to directly. Both shapes updated in the same action. Nothing pushed back in.

The store owns the change handler, so React Flow's report *is* the update — there is no second write to echo back: ```ts import { create } from "zustand"; import { applyNodeChanges, type Node, type NodeChange } from "@xyflow/react"; type EditingStore = { nodes: Node[]; edges: Edge[]; // React Flow calls this directly. No effect, no sync, no guard flag. onNodesChange: (changes: NodeChange[]) => void; updateNodeConfig: (id: string, patch: Partial) => void; }; export const useEditingStore = create((set) => ({ nodes: [], edges: [], onNodesChange: (changes) => set((state) => ({ nodes: applyNodeChanges(changes, state.nodes) })), updateNodeConfig: (id, patch) => set((state) => ({ nodes: state.nodes.map((node) => node.id === id ? { ...node, data: { ...node.data, config: { ...node.data.config, ...patch } } } : node, ), })), })); ``` And a node component subscribes to its own slice, not to the array: ```tsx function DynamicNode({ id }: NodeProps) { // Re-renders when THIS node's config changes. Not when node 12 moves. const config = useEditingStore((s) => s.nodes.find((n) => n.id === id)?.data.config); const status = useExecutionStore((s) => s.statusByNode[id]); const updateNodeConfig = useEditingStore((s) => s.updateNodeConfig); return updateNodeConfig(id, p)} />; } ``` That second line is the whole point of the migration. Node 12's status arriving from a poll touches node 12's component. Nodes 1 through 11 do not hear about it — they never subscribed to `statusByNode` as a whole, only to their own key. The rest fell out of it: - nothing is pushed back into React Flow, so there is no loop - no loop, so no guard flag and no hash - the 300ms debounce went away, because there was nothing left to protect - when we need React Flow's own internals, `useStore` and `useStoreApi` are right there — the same store, not a black box behind a Context ## Do the two stores talk to each other? They have to. A canvas node shows its live execution status right on the node — running, done, failed, and eventually its output. So the canvas reads from both.
EDITING STORE EXECUTION STORE nodes + edges positions · viewport · zoom node config per node selection status per node pending · running · done output per node kept fresh by polling one canvas node its config its status
Separate stores, read together at the leaf. Each node subscribes to its own key in each store, not to either store as a whole.
The catch that shaped this design: **the user can keep editing — even delete — a node while it is still executing.** That is exactly why execution is not folded into the editing store. - The editing store stays focused on what the workflow *is*: nodes, edges, configs, positions. - The execution store tracks what is *happening*: status, progress, output, kept in sync by polling. - If a node is deleted mid-execution, the editing store drops it immediately. The execution store keeps polling in the background and quietly ignores results for nodes that no longer exist. If those lived together, deleting a running node would mean untangling one from the other on the hot path. They do not, so it does not.
Outcome

The lag at 30–40 nodes went away. A drag settles in a frame instead of 350ms, one node moving costs 1 re-render instead of 20–50, and about 150 lines of sync, guard-flag and hashing code were deleted rather than rewritten. Editing and execution now run side by side without either dragging the other down.

## What I would tell someone before they start Context is not bad. It is excellent for state that rarely changes — theme, auth, locale, a feature flag. It is a dependency-injection tool that happens to hold state. For state that updates often and is read by many components — a canvas with dozens of interactive nodes — you need granular subscriptions from day one, and no amount of provider-splitting or `memo` will retrofit them. And the part I would say first: **if the library you build on has already made a state decision, the cheapest path is usually to make the same one.** We did not choose Zustand on its merits alone. We chose to stop fighting a store that was already running underneath our canvas. The wider version of this story — the workflow platform, the execution engine, the QC that stopped humans reviewing ten thousand images — is in [Foosh: From One Prompt Box to Ten Thousand Rows](/blog/foosh-engineering-experience). ## Common questions ### Why use Zustand with React Flow instead of React Context? Because React Flow already stores its own graph state in Zustand. Node positions, viewport, drag state, selection and connections all live in a Zustand store inside the library. If you keep your node data in React Context as well, two systems own the same data in two different shapes and you have to write a transform and a sync layer between them. Putting your state in Zustand removes that layer: React Flow's change handler becomes an action on your own store, both shapes are updated in the same call, and components subscribe per node instead of per provider. ### Why does a React Flow canvas get slow when you add more nodes? Usually it is not React Flow, it is how the node data is subscribed to. If every custom node reads from one React Context, or from the whole nodes array, then changing one node re-renders all of them plus the sidebar and any panel reading the same value. On our AI workflow builder that showed up at around 30 to 40 nodes as sticky dragging and stuttering zoom, and profiling showed 20 to 50 re-renders for a single node being moved. The fix is granular subscriptions, so a component only re-renders when the specific slice it reads actually changes. ### How do you update one node's data in React Flow without re-rendering every node? Keep the per-node data in a Zustand store keyed by node id, and have each custom node component select only its own key. The default approach — call setNodes on the whole array to change one node's data — makes React Flow reconcile the entire array, and any node component reading that array re-renders with it. With a store, updating node 12's config or execution status touches node 12's component only; nodes 1 to 11 never subscribed to it. ### Is React Context bad for state that changes often? It is the wrong tool for it. React Context has no concept of a partial subscription: any component that consumes a context re-renders whenever any value in that context changes, even if it only reads one field of it. That is fine for state that rarely changes — theme, auth, locale, feature flags — and it falls apart for state written many times a second, like a canvas being dragged. For that you need a store with selector-based subscriptions, such as Zustand, Redux with React-Redux, or Jotai. ### Does splitting React Context into smaller providers fix re-render problems? Not on its own. We tried it before switching libraries — one provider per node, with memoised consumers — and changing one node still re-rendered the other nodes and their internals. Slicing providers thinner does not change the rule that everything reading a provider re-renders when that provider's value changes, and React.memo cannot help when the value a component is subscribed to is a new object on every update. It is worth trying because it is cheap, but do not plan around it working. ### Zustand or Redux for a node editor or canvas app? Both can solve the re-render problem, because both support selector-based subscriptions. We chose Zustand for three reasons: creating a store is a few lines with no provider or reducer wiring, selector subscriptions are the default rather than something you have to be disciplined about, and React Flow's own internal store is Zustand, so there is one state model in the app instead of two. Redux earns its boilerplate on larger apps with many teams and strict action logging; for fixing a canvas re-render storm it was a slower, heavier route to the same place. ### Should editing state and execution state be in the same store? Keep them separate if the user can edit while something is running. We use one Zustand store for editing — nodes, edges, positions, configs — and a second for execution — status, progress and output per node, kept fresh by polling. The reason is a product requirement: you can edit, and even delete, a node while it is still executing. With separate stores the editor drops a deleted node immediately and the execution store just ignores results for nodes that no longer exist. In one store, every polled update would be writing to the object the editor is writing to. ### What state does React Flow keep internally, and can you read it? React Flow keeps node positions, the viewport and zoom, drag state, selection and in-progress connections in an internal Zustand store. You can read it directly with its useStore hook for reactive access, or useStoreApi when you need the current value without subscribing. That is one of the practical advantages of using Zustand for your own state too: the library's state and your state are the same kind of thing, rather than one being a black box behind a Context you cannot select into. ### How many nodes can React Flow handle before performance becomes a problem? There is no fixed number, because the ceiling is set by how your node components subscribe to data rather than by the library. With every node reading one React Context we felt it at 30 to 40 nodes, especially with an execution running at the same time. After moving the same graph to Zustand with per-node selectors, that lag disappeared and the 300ms position debounce we had added to hide it was deleted. Fix subscriptions first before reaching for virtualisation or node culling. --- ## Foosh: From One Prompt Box to Ten Thousand Rows Published: 2026-08-27 · Category: experience · URL: https://www.mukeshbishnoi.com/blog/foosh-engineering-experience Building an AI workflow platform end to end. Why React Context could not keep up with React Flow, how run state moved from memory to DynamoDB to S3, why we started at five parallel rows, and what happened when an LLM started doing the QA. Foosh is for companies that would otherwise book a studio. You want a campaign for your product. The old way is a photoshoot. Studio, photographer, lights, model, retouching. Weeks of it. Next quarter you need a new campaign, so you do it again. We build the other way. Upload the product once. Build a workflow. Generate the campaign. The output goes to the website, to ads, to banners, and to video. I joined in April 2025 as the founding engineer. This is what the year looked like.
53 models wired in
18 node types
10,000 rows in one campaign
240 frames per video segment
## Where I started When I joined, the product was about LoRA training. The problem it solved: - A company has a product. A shoe, a bottle. - A general image model has never seen that exact bottle, so it cannot draw it. - So you train it. Upload photos of the SKU, train a LoRA on top of Flux. - Now the model knows your object, and you can put it in any scene. I was doing frontend only then. Training is slow. It is not a request that returns. It runs for a long time, and the user is sitting there wondering what is happening. So the model library page was really a status page. Which LoRAs are training, what stage each one is at, which ones are done. The whole screen existed to answer one question. Is my model ready yet. ## Why the product changed Two things happened together. The models got better. Newer image models take a few reference images at generation time and give you your product back, correctly. No training run. No waiting. And the business needed more than one feature. Training a LoRA per object is one thing. Customers wanted campaigns. Many images, many products, many formats, and video. So the product moved. From "train a model on your object" to "build a workflow that makes the whole campaign". ## Where I stopped being frontend only I built the workflow builder frontend first. Canvas, nodes, edges, the editor. Once it was running we had about five models wired in. Then the bottleneck moved. Every new model was backend work. Every new node was backend work. Waiting for that half was slower than learning it. So I went fullstack. After that I was adding models and nodes end to end, canvas to Lambda. ## The app: three surfaces The first thing I built on the new platform: - **Image library.** Every image the workspace has generated. - **Generation.** Prompt in, generation showing as it runs. - **Model library.** LoRAs training, and the trained ones. Each one had its own problem. ### The library got heavy A real workspace holds thousands of images. Not hundreds. AI generated images are not small files. What went wrong was ordinary: - scrolling was slow - images popped in late, after you had scrolled past - the tab kept eating memory the longer you stayed Two fixes. Neither is clever. **Load a page at a time.** The gallery pulls 20 at a time and fetches the next page when you are about 1000 pixels from the bottom. You never hold the whole workspace in the page. **Stop sending full size images to a grid.** This was the bigger one. We put Gumlet in front of our S3 bucket. It returns the same image at whatever size we ask for, cached per size. The grid gets a 300px thumbnail at around 50KB, not a 4MB original that the browser then shrinks into a small box. Four variants of every image: | Variant | From | Used for | | --- | --- | --- | | thumbnail | CDN, w=300 | the grid, ~50KB | | medium | CDN | preview | | full | CDN, quality=100 | viewing one image | | original | S3 directly | download | The last row cost us a bug. The CDN compresses whether you ask it to or not. So a download link pointing at the CDN gives the user a compressed file that they think is the original. Viewing goes through the CDN. Downloading does not. ### We poll, and that was on purpose Fire a generation, get an execution id, then keep asking if it is done. We have polled since the beginning. We did not try websockets and give up. - On Lambda, holding a socket open means paying for a function that sits there doing nothing. - You also have to keep that connection alive across a system built to be stateless. - The job takes thirty seconds to five minutes. Nobody can tell the difference. The generation is slow because the model is slow. A socket would not have made it faster. ### Every component was calling axios This was the mess I cleaned up. Every component made its own axios call. So there was no caching anywhere. Open a screen, it fetches. Go back, it fetches again. Open a campaign from last month, one that is finished and cannot change, and it fetches that too. That last case is the point. Most of what this app shows is finished work. Asking the server for it every time is a slower page for no reason. So it all moved to React Query, in one place. One axios instance, one query client, one set of constants, and service files split by domain. Components call a hook. They do not know axios exists. What we got: - finished things stay cached - live things stay fresh - when something does change, you invalidate one key instead of hunting for every component that fetched it ### Clerk, and the workspace mapping Auth is Clerk. The part worth saying is that we did not build our own idea of a team on top of it. A Clerk organisation is the workspace. The org id is the workspace id. Credits, workflows, gallery, members and roles all hang off it. There is no separate workspace table with its own ids to keep in sync. That kind of mapping drifts. Roles are viewer, member, admin, owner. They are enforced on the backend, not hidden in the UI. A viewer cannot run a workflow, and that check happens where credits get deducted. ## The workflow builder This is the piece I spent the most time on. Instead of one prompt box, give people a canvas. Drop nodes, wire them together. Each node is a model call or a piece of logic. A real workflow looks like this: ``` product image → enhance the prompt with an LLM → generate 4 variations → upscale the good one → drop it into a banner template ``` Save it, and run it again next month with a different product. It runs on React Flow. Today there are 18 node types and 53 models behind them, across nine providers. ### One node component, not eighteen There is no ImageGenerationNode component. There is no VideoGenerationNode component. There is one DynamicNode, and it reads what to render from config. - The backend has one config file describing every node type. Its inputs, its parameters, its sockets, its outputs. - The frontend fetches that config and builds the node from it. - Adding a model is a config change. Nobody writes a React component for it. ### Two contexts State started in React Context, and we had two on purpose: - **the save flow** — the graph itself. Nodes, edges, positions, parameters. The thing that gets persisted. - **the execution flow** — what is happening now. Which node is running, what came back, what failed. Splitting them made sense. They change at different rates. The graph changes when you edit it. The execution state changes every time you poll. That part was fine. The problem was somewhere else. ### The real problem was two owners React Flow already keeps the graph in its own internal store. It has to. It is the one drawing the canvas and handling the drag. So the moment we also kept the graph in context, two things owned the same data, in two different shapes: - our shape: model, parameters, sockets, outputs - React Flow's shape: position, type, data Something has to transform between them. That transform sat in the middle: ``` Context → Transform → Sync → ReactFlow ``` Now drag one node: ``` React Flow updates its own store, tells us → we update our state → the transform runs, produces a new nodes array → we push that array into React Flow → React Flow sees a new array, treats it as a change, tells us → repeat ``` That is a loop. To stop it you add: - a guard flag, so we can say "this update came from me, ignore it" - a hash check, because you need to know if the incoming array is actually different or just a new array with the same contents So on every render we were serialising the entire nodes array and comparing it. Twice. ### What the lag looked like Dragging fires a change event on every mouse move. Sixty times a second. You cannot rewrite your state, run the transform and hash the whole graph sixty times a second. So we debounced position updates by 300ms. That debounce is what you could see: - React Flow paints the node under your cursor straight away, because it owns the canvas - our state was a third of a second behind - so while you dragged, anything reading our state saw the old position. The inspector, connected nodes, validation - drop the node and everything snapped into place Measured: about 350ms for a drag to settle, and 20 to 50 re-renders for one node moving. That is what context does. Change the value and every consumer re-renders, and every node on the canvas was a consumer. Around 150 lines existed only to keep the two copies agreeing. ### The fix was to stop being the second owner We moved to Zustand. Not because Zustand is fashionable. Not because context is bad. Because React Flow's internal store is already Zustand. That is the whole reason. We were not picking a state library. We were choosing to stop fighting one that was already there. ``` before: Context → Transform → Sync → ReactFlow after: Zustand store ═══════════════ ReactFlow ``` A drag now goes: React Flow reports the change to a handler that belongs to our store, the store updates both shapes in the same action, and only the components subscribed to that node re-render. - Nothing gets pushed back in, so there is no loop - No loop, so no guard flag and no hash - Zustand subscribes per selector, so a component asking for one node re-renders when that node changes, not when anything moves - The debounce went away, because there was nothing left to protect If the library you build on has already made a state decision, the cheapest path is usually to make the same one. I wrote the long version of this one up separately — the Context code that broke, why splitting the providers and adding `memo` did not save it, why Redux lost to Zustand, and how the two stores talk to each other: [Why We Use Zustand with React Flow to Manage State](/blog/context-to-zustand-migration). ## Nobody checks ten thousand images by hand Once campaigns arrived, one run could produce thousands of outputs. Each one still needs a human decision. Good enough to publish, or not. That does not scale. So there are two layers of QA, and the second only works because of the first. ### The QC node QC is a node you drop in after your generation nodes. It takes the output, optional reference images, and your criteria. An LLM looks at each output and says accepted or rejected, with a reason. Two details matter more than they look. **It does not filter.** It does not throw away rejected outputs. It attaches its verdict and passes everything through. A model that silently deletes your work is not something you want in a pipeline. It labels. You decide. **The criteria are just text.** The customer writes what they want in plain English. The label must be readable. The bottle must match the reference. The logo must not be cropped. We do not make them build rules in a rule builder. "Does this image look right" is not something you can write as a rule. That is why an LLM is doing the job. We have our own QC rules underneath, and when the node runs we take what the user typed and turn it into the rules for that run. It runs on Gemini 3 Flash, with GPT-5.2 as fallback if Gemini fails three times. ### Keeping the verdicts small Two problems here, both only visible at bulk scale. **The verdicts are too big to store inline.** Every output gets a verdict and a written reason. Across a campaign that is a lot of text inside a database record that has a size limit. So the full result goes to S3, and DynamoDB keeps only the key. About 50 bytes per item instead of the whole thing. **The verdicts have to land on the right image.** If a node made four images and QC evaluated four images, verdict three belongs to image three. Exactly, not roughly. So indexes are preserved the whole way through, and reason ids are typed per media so an image verdict and a video verdict cannot collide. That sounds obvious. It is the thing that quietly breaks when arrays get flattened and re-aggregated across a distributed run. ### Then the human layer After the AI labels everything, the customer reviews. - They can go through every output. - Or filter to only what the AI accepted. - Each output can be approved or rejected with a reason. - The download can be limited to approved ones. The reviewer never starts from nothing. The AI's verdict and reasoning are already on each output. So a human is confirming or overruling a first pass, not forming an opinion from scratch on image four thousand. ### What actually happened This is the real result, and I did not expect it. It goes in stages: 1. At first the customer checks everything. They do not trust it yet. 2. Then they look only at what the AI accepted, because that is where the usable images are. 3. Then, once their rules have held up run after run, they stop opening the folder. The QC node did not save time by reviewing faster. It saved time by letting people stop reviewing. That only happened because the node labels instead of deletes. Being wrong stayed visible, so trust could be earned instead of assumed. ## Giving the canvas to people who do not want it The canvas is good if you like building things. Most people do not. They want their campaign. So a workflow can be published as an app. - When you publish, you choose which inputs the user fills in and which outputs they get back. - Those two choices become a form. - The user sees the form. Product image, tagline, press go, get the banner. - No nodes. No wiring. No idea a graph exists behind it. That user is sometimes the customer, and sometimes a colleague of the person who built the workflow. The marketing person who is never going to learn a node editor and should not have to. The builder stays for people who build. The app is for everyone else. ### Every published agent is a public page A published app gets its own public URL, and that page has to work for someone who has never logged in. Which means it has to be a real page, not an empty shell that fills itself in after the JavaScript loads. This is why the frontend is Next.js. - The page is rendered on the server, so the agent's name and description are in the HTML. - Each agent gets its own title, description, and Open Graph and Twitter cards, built from that agent's own data at request time. - Those responses are cached for a few minutes, so a popular agent is not hitting the API on every visit. The practical result: paste an agent link into WhatsApp or Slack and the preview shows that agent, with its own thumbnail and description. Not a generic card with the company logo. Same for shared workflows. Each share link is its own page with its own metadata. Every agent someone publishes is a page that can be found, linked and previewed on its own. ## Then somebody wanted five hundred at once One at a time is a demo. The real job is a catalogue. So, campaigns. Upload a CSV, one row per thing you want made. Each row names the app to run and the values to fill in. Up to ten thousand rows. That is where this stopped being a web app and became a queue. ### Twenty at a time, and we started at five Rows run in parallel, but not all at once. Today it is twenty. We started at five. The limit was not the models and it was not cost. It was Lambda concurrency. - Every row is a workflow. - Every node in that workflow is a Lambda invocation. - A few hundred rows at once becomes a very large number of concurrent functions, very fast. We were tripping the account limit and taking down everything else with it, including workflows that had nothing to do with the campaign. So we kept it low and moved it up as we learned the real ceiling. Five, then more, now twenty, with Step Functions Distributed Map fanning out and keeping concurrency capped. Nothing clever about the number. The clever part was accepting that a slow safe setting beats a fast one that takes the platform down. ### Validate before you charge anyone The CSV is checked before anything is charged and before a single row runs. - Do the named apps exist in this workspace - Are the required columns there - Are they filled in - If an app name is close but not exact, we say which one we think they meant The reason is simple. A ten thousand row campaign that dies on row four hundred because of a typo in a header has already spent real money. Credits work like a hotel deposit: - We take the full estimate up front as one atomic reservation. - If the workspace cannot cover it, the campaign never starts. Nothing is half charged. - At the end, whatever was not used is refunded against the original transaction. Rows fail. A model times out, a provider errors. Those rows get a record straight away with the actual error on it. Then you rerun that row, not the campaign. ### Where virtualisation actually mattered A finished campaign has thousands of images. The reviewer scrolls a filmstrip through them. This is where we needed real virtualisation. Only the rows in view exist in the DOM. The rest is space. The gallery gets away with paging 20 at a time because you scroll it slowly and look at things. A review filmstrip is different. People drag through it fast, all the way, repeatedly. Paging is not enough when someone is scrubbing. Two different scrolling problems, two different answers. Worth saying, because "just virtualise everything" is the usual advice and it is not free. ### Downloads had to leave Lambda Zipping thousands of full size images does not fit in a Lambda. It is not a memory problem, it is a time problem. Lambda has a hard ceiling on how long it may run and a big download goes past it. So downloads run on ECS Fargate, where nothing times out. - The zip is split into parts at about 1.5GB - A folder is never split across two parts. Unzipping half a folder from part one and half from part three is miserable - Finished downloads are cached, but the cache checks whether the campaign was touched after the zip was built. Rerun a row and the old zip is no longer the truth ## What actually runs a workflow A workflow is a graph, and a graph is not a shape a web request handles. Some nodes wait on others. Some run side by side. One node can take five minutes because a video model is thinking. So execution runs on AWS Step Functions, with Lambda doing the work. There is one state machine, not one per workflow. We do not generate a machine every time somebody saves a canvas. The machine is generic: ``` read the graph → work out what can run now → prepare each node's inputs → invoke the right Lambda → write the output back → repeat ``` The graph is data. The machine that walks it stays the same. ### The run state moved twice This changed the most, and each move was forced. **Stage one: in memory.** Execution state travelled through the run. Each step handed its outputs to the next. Simple, no storage, nothing to clean up. It fell over as soon as workflows got real. Outputs are not small. A text node returns paragraphs. An image node returns URLs and metadata. Chain six nodes and you are carrying all of it along. Everything had to stay small enough to pass along, and it stopped being small. **Stage two: DynamoDB.** So run state moved into a table. Each node writes its output to the execution record. Anything that needs it reads from there. - Nothing is carried around any more - The frontend polls the same record - A run that dies leaves its state behind, so you can see how far it got That worked for a long time. Then the next wall: a DynamoDB item cannot be bigger than 400KB. A workflow with a lot of nodes, each holding its outputs, passes 400KB more easily than you would think. And the failure is ugly. The run does not slow down. It refuses to write. **Stage three: S3, with a pointer.** Now, when node data is too big for the item, it goes to S3. What stays in DynamoDB is a small reference saying this lives in S3, here is the bucket and key. The important half: every path that reads node data resolves that pointer first. The part that prepares inputs, the part that finalises the run, the part that answers the poll. All of them ask "is this a pointer" before using it. Nothing downstream needs to know. The frontend has never seen an S3 reference. Three stages, same pattern each time. State starts wherever is easiest, then moves outward as it outgrows the container it is in. Memory, then a row, then a bucket with a pointer where the row used to be. ### Two nodes finishing at the same moment Parallel execution brought its own bug. If two nodes finish together and both write to the same execution record, one can quietly overwrite the other. Nothing errors. You get a workflow that reports success with one node's output missing. So the record carries a version. You write on the condition that the version is still what you read. If someone got there first, re-read and try again. Ordinary optimistic locking. Worth mentioning because this bug is invisible until a customer asks why an output is missing from a run that said it succeeded. ### Running one node Small thing, big difference. You are building a workflow. You tweak a prompt on node six. You want to see what node six does now. You should not have to re-run nodes one to five. They have not changed, and they cost money. So you can run a single node. Everything else keeps the outputs it had. Its inputs come from the nodes already feeding into it. One rule: you cannot run a single node while the whole workflow is running, because both would write to the same place. Two separate nodes at once is fine. ## The video pipeline One client needed something the platform did not do. Take videos their users had already recorded, and put them somewhere else. Same person, same performance, new background. Almost every decision in it comes from one constraint. ### Everything is shaped by 240 frames The model doing the background replacement will not take a video longer than 240 frames. At normal frame rates, about ten seconds. Real videos are minutes long. So step one is not clever, it is forced. Split the video into segments that fit under the limit. Everything after that exists because we cut the video up and have to put it back together. ### Each segment needs to know the new background The replacement model needs a reference image showing the scene we want. So we generate those, one per segment, using Gemini and Nano Banana. Take a frame from the segment, compose the person into the chosen setting, and that image becomes the instruction for what the segment should look like. Then each segment goes through background replacement on its own, in order. ### Putting it back without the jump This took the most fiddling, and not for the reason you would expect. Cut a video into ten second pieces, process each separately, join them end to end, and you can see every joint. Each segment was processed on its own, so lighting or framing lands slightly differently. The eye catches every one of those small shifts. The fix is to overlap segments by 60 frames and fade across the overlap. The shift is still there. You just cannot see it, because it has been spread over two and a half seconds instead of landing between two frames. Audio is cut at the midpoint of the same overlap, so one segment hands over to the next without a gap and without both playing at once. Captions go on at the end, transcribed with Sarvam. ### It has to survive being interrupted A long video becomes a lot of segments. Each takes real time. A full run is long. So the pipeline keeps a progress file and marks each segment done as it finishes. If it dies in the middle, and over a run that long it will, starting again picks up from the first segment that is not done. Any pipeline that runs for hours needs this. It is always tempting to skip until the first time you lose an hour of work. ## Two smaller things ### Model output has to be data, not prose A lot of nodes call an LLM. An LLM left alone writes you a paragraph. That is useless in a workflow. The output is not read by a person. It is saved as the node's output and wired into the next node. Something has to pick it up, store it, and hand the right piece to whatever comes next. You cannot do that with a sentence. If a node returns "Sure! Here are three taglines for your product: ...", there is nothing to grab. Which part is the tagline. Are there three. Where does one end. So nodes ask for structured output. A defined shape, every time. The failure mode without it is the real argument: - the node does not fail - it returns its paragraph and reports success - the break shows up two nodes later, where something used a field that was never there The error appears far away from the thing that caused it. In a graph, that is the worst kind of bug to chase. Structure at the boundary means a bad node fails at itself. ### Bring the assets from where they already are Products come with photos, and every product needs several. Front, back, angles, packaging. The entity library is where those live. An entity is a product or a brand. A name, and the images that belong to it. A workflow then refers to it by name. You do not attach files to a workflow. You say which product this run is about. The upload is the part worth mentioning. In this industry everybody's product photos are already in Google Drive, in folders, one folder per product. That is just how it is done. So we read the Drive folders directly. Point us at the parent folder, we iterate through it, and each folder becomes an entity with that folder's name and the images inside. The alternative is what they were doing before. Download everything from Drive to your laptop, then upload into our system one file at a time. For a few hundred products that is a day of somebody's life, spent doing nothing. The feature is not bulk upload. The feature is not making somebody restructure their filing to use your product. ## What I take away Two things. **Almost none of the hard problems were AI problems.** Calling the models was the easy part. The hard parts were ordinary engineering. Who owns the state. Where it lives when it outgrows the row. How many things can run before the account falls over. Whether anyone trusts the output. **The limits designed the system.** Very little here came from a whiteboard. - 240 frames is why the video is split. - 400KB is why run state went to S3. - Lambda concurrency is why we started at five parallel rows. - React Flow owning its own store is why we moved to Zustand. Every one of those is a ceiling somebody else set. The work was noticing the ceiling early and building with it instead of against it. I started this year doing frontend. I finish it having built both halves, mostly because waiting for the other half was slower than learning it. ## Common questions ### How do you run a node-graph workflow on AWS? We use one generic AWS Step Functions state machine with Lambda doing the work, rather than generating a state machine per workflow. The machine reads the graph, works out which nodes can run now, prepares each node's inputs, invokes the right Lambda, writes the output back and repeats. The graph is data; the machine that walks it never changes, so saving a new canvas does not deploy anything. ### Why does a DynamoDB write fail on a large workflow run, and how do you fix it? A DynamoDB item cannot exceed 400KB, and a run record holding every node's output passes that more easily than you would expect. The failure is abrupt — the run does not slow down, it refuses to write. The fix is to put oversized node data in S3 and keep a small pointer in the item, then make every path that reads node data resolve that pointer first, so nothing downstream needs to know where the data actually lives. ### Why would two parallel nodes lose each other's output in a workflow run? Because both write to the same execution record and one silently overwrites the other. Nothing errors, and the run reports success with an output missing. The fix is ordinary optimistic locking: keep a version on the record and write on the condition that the version is still what you read, re-reading and retrying if someone got there first. ### Should you use polling or websockets for long-running AI generations? We have polled from the beginning and would again on this architecture. On Lambda, holding a socket open means paying for a function that sits doing nothing, and keeping that connection alive across a deliberately stateless system is work. Generations take thirty seconds to five minutes, so nobody can tell the difference between a socket and a poll — the model is what is slow, and a socket would not make it faster. ### How do you make an image gallery with thousands of AI-generated images fast? Two ordinary fixes did it. Page the data — we fetch 20 at a time and load the next page about 1000 pixels from the bottom, so the whole workspace is never in the page. And stop sending full-size images to a grid: a CDN in front of the bucket returns a 300px thumbnail at around 50KB instead of a 4MB original the browser then shrinks. One warning — a CDN compresses whether you ask it to or not, so downloads have to bypass it or users get a compressed file they think is the original. --- ## A Year and a Half at Unolo: Four Problems Worth Writing Down Published: 2026-08-24 · Category: experience · URL: https://www.mukeshbishnoi.com/blog/unolo-engineering-experience Churn caught a month early, a landing page cut from 2.5s to 0.5s, a dashboard the customer assembles themselves, and a map that went from crashing at 4,000 points to smooth at 100,000. Unolo is a field force management platform. Companies use it to track the field employees who leave the office every morning — salespeople, on-ground staff — and to follow the work they're sent out to do. The customer base is enterprise: large organisations with distributed teams, including names like **Tata** and **Uber**. Publicly, Unolo states it's used by 1200+ companies and 50,000+ field staff. I was a **frontend engineer** there from **November 2023 to April 2025**, inside an engineering team of about **30 people**. Four pieces of work from that time are worth writing down properly: an internal usage-tracking platform that changed when the sales team learned about churn, a performance overhaul of the public site, a configuration-driven dashboard I led, and a rewrite of the map system that took it from choking at 4,000 points to comfortably rendering 100,000.
10–20% churn reduction
2.5s → 0.5s landing page load
4k → 100k points on the map
4 engineers led on the dashboard
## The stack Everything below happened on this stack: | Layer | What we used | | --- | --- | | Application & marketing site | React.js, Next.js | | State | Redux | | UI & styling | Material UI, Tailwind CSS | | Dashboard surface | react-grid-layout (drag & resize) | | Charts | Recharts | | Maps | Google Maps, with deck.gl layered on top for custom layers | --- ## Catching churn before the cancellation ### The problem The sales team used to find out about a churning account at the worst possible moment: when the customer cancelled the subscription. By then there was nothing left to do. The signal arrived after the decision had already been made. The premise of the project was simple. A customer about to churn almost always shows it in their usage first. If we could surface that early, someone from our side could reach out *while the account was still saveable*, ask what was going wrong, and fix it. So we built an internal platform for the sales team whose only job was to make client usage visible. ### What "usage" actually means here Usage was never one number. It was a set of signals about how alive an account was: - how many **active users** the client has - how many users are **logging in daily** - how many **tasks their teams are creating** - how many **tasks are being assigned** - volumes across the product's core entities — **tasks, sites, companies and visits** And on top of every one of those, **comparison over time** — because the absolute number tells you far less than the direction of travel: last week vs this week, yesterday vs today, last month vs this month, last year vs this year. > A client with 400 active users is healthy. A client that had 400 last month > and has 180 this month is a phone call that needs to happen today. The > comparisons were the product. ### How it was built The **frontend was a completely separate project** — its own application, not a section bolted onto the customer-facing dashboard. The audience was internal, the access rules were different, and the views had nothing to do with what a customer sees. The **backend was the product's actual backend**. We didn't stand up a parallel data store; we added a set of **additional APIs on the existing backend** that exposed the usage information in the shape this platform needed. ### What was hard This project was data-heavy, and that's where most of the engineering effort went: - **A lot of filters.** Every view needed to be sliceable — by client, by entity, by date range, by comparison window. The filter surface itself became a significant piece of state to manage. - **A lot of conditional logic on the frontend.** The branches for handling combinations of filters, comparison modes and data shapes grew fast, and keeping that readable was a constant fight. - **A lot of data manipulation on the client.** Reshaping, grouping, deriving the comparison deltas — much of that transformation happened frontend-side. - **Visualisation mattered as much as the data.** A raw table of usage numbers does not make a salesperson pick up the phone. A visible drop does. Getting the presentation right was part of the requirement, not polish on top of it.
Outcome

The lead time changed completely. Instead of learning about a problem at cancellation, the sales team was seeing trouble roughly a month earlier — early enough to call, ask what the customer was struggling with, and in a lot of cases simply get told, and fix it. Churn came down by roughly 10–20%.

--- ## Making the landing site fast ### Why we picked this up Three signals landed at roughly the same time, and they were all the same problem wearing different clothes: - **PageSpeed** scores on the landing site were poor. - **SEO** was suffering — the page took so long to load that **indexing was being affected**. - **Marketing complained** about how slowly the page came up for visitors. For a marketing site that's the whole ballgame. The landing page is the first thing a prospective client sees, and if it takes several seconds to paint, a good number of them never see it at all. An unindexed page might as well not exist. ### How I found what was wrong I profiled with **Lighthouse** and **Chrome DevTools**, deliberately running under **different throttling profiles** rather than on a fast office connection — because the visitor on a mid-range phone on a mobile network is the visitor whose experience is actually broken. The audit pointed at two things consistently: **images** were taking a large share of load time, and **fetching CSS and JS** was the other large share. Everything below follows from those two findings. ### Image optimization Images were the single heaviest thing on the page, so this was the first pass: - **Reduced the image sizes** outright — the originals were far larger than anything the page actually needed. - **Converted to WebP**, with compression on top. - Introduced **`srcset` with proper dimensions per screen size**, so a phone downloads a phone-sized image instead of a desktop-sized one scaled down in the browser. - Moved images onto the **CDN**, so they were served from close to the visitor. ### Chunk size reduction The CSS and JS together were bloating the page. Two moves: - **Removed unused CSS** that was being shipped in the bundle but never applied. - **Chunked and split the bundle by page and by section**, so the browser fetches what the **visible section** needs first and the rest arrives afterwards. > The browser should not have to download the code for the bottom of the page > before it can paint the top of it. ### FCP and LCP The two problems above were exactly the two things holding these metrics back. The **heavy images** were causing significant layout shift and painting problems, and it got noticeably worse across responsive breakpoints, where the wrong-sized image was loaded and then reflowed. The **bundle size** was the other half — once there was less to paint and less to fetch before painting, the numbers moved. Fixing images and bundle size *was* the LCP/FCP work. They weren't separate tasks. ### Lazy imports for components Every page on the site is made of a lot of sections and components, and very few of them are on screen when the page first loads. So the rule was simple: **visible components get brought in first; components outside the visible region load after** — via lazy imports, rather than being bundled into the initial payload. ### Assets on the CDN We had **clients across different regions**, and the site carried a lot of assets plus large CSS and JS files. Serving all of that from a single origin meant how fast the site felt depended largely on how far away you happened to be. Putting the assets on a **CDN** removed that. Alongside it, going through the CSS and JS carefully and **tree-shaking out what wasn't used** brought file sizes down substantially — and a smaller file served from a nearer edge is a compounding win, not an additive one.
Outcome

Load time for the images and the first page went from around 2–2.5 seconds down to roughly half a second to one second. That came from three things working together, not one: image sizes came down, whatever is visible is what gets shown first, and assets that aren't required at that moment are not loaded at all.

--- ## A dashboard the customer assembles themselves *A project I led — one frontend (me), two backend, one QA.* ### The problem with a fixed dashboard The dashboard was **fixed**. Everyone got the same screen. That works right up until you look at who's actually using it. A single customer has **many different admins — different teams, different regions**. The admin for one region needs that region's data; the admin for another team cares about something else entirely. The data set behind the product is large, and not all of it is relevant to everyone. The real insight was that this is not a permissions problem, it's a **priority** problem. Two admins can both be entitled to see the same numbers and still want completely different screens, because their field of interest differs based on what they're responsible for. A fixed dashboard forces both of them to look past most of the screen to find their one number. So the dashboard needed to stop being a screen we design, and start being a screen the user assembles. ### The design: JSON describes the dashboard The whole dashboard is driven by a **JSON configuration**, which carries everything that makes a dashboard a dashboard: - **what position** a component sits at - **what data** it's bound to - **how it looks** - **how that data should be represented** Nothing about a given dashboard is hardcoded in the UI. The UI is a renderer; the JSON is the dashboard. ### How a config entry becomes a component Rendering works through **component mapping driven by type**. The type tells the renderer what it's dealing with at two levels: 1. **What kind of component is this** — a data component, or a visual one? 2. **If visual, which visual** — a graph, and then *which* graph, or a comparison view. That two-level split is what keeps the system extensible. Adding a new chart type doesn't mean touching the renderer's logic; it means registering another entry in the map. ### Position, layout and the component library **The complete layout and position is configurable.** Components can be arranged, moved and sized, and that arrangement is part of the configuration — which is what makes the dashboard genuinely *per admin*, rather than a fixed grid with swappable contents. The component library covered the full range: **all chart types**, **tables**, **widgets** carrying data, and **comparison views**. And crucially — **each component pulls its data from a different API, and which API it calls is configured in that component's JSON**. The data binding is part of the config, not part of the component. A chart doesn't know where its numbers come from; the config tells it. ### Lazy importing and the component API Two things kept this from turning into a bundle problem: - **Lazy importing**, so the application only loads the components a given dashboard actually uses — not every chart type in the library on the chance that someone might place one. - A **compound component API** — components exposed in a nested form such as `` and `` — which kept the surface well balanced: readable and grouped for the developer, while still allowing each variant to be pulled in only when needed. ### What was actually hard Not the rendering. It was **the data flow and the API call flow inside every component**. When every component fetches from its own configured endpoint, you no longer have one page making a coordinated set of requests. You have N independent components, each deciding on its own what to call, when to call it, what to do while it waits, and how to behave when its call fails — all while the user can add, remove, move and resize them. Getting that flow right, consistently, across every component type, was where the real work sat.
Outcome

A new dashboard stopped being an engineering task. It became clicks. Anyone can choose what they want to see and what they don't — arrange it, bind it, and it's there. No waiting on a build cycle, no timeline, no ticket.

The shift in what the product asks of the user is worth stating plainly. **Before**, whether you found it useful, whether it mattered to your team, whether it was something you'd want to share with another admin — none of that was considered. The message was *this is the information, this is the order, look at it.* **After**, there's real control over what each admin sees, over the visibility of data according to that admin's area of administration, and over which data is most valuable to which administrator so that it appears first. And they can adjust it themselves. What used to be a single fixed arrangement became a screen that reflects the person looking at it. --- ## The map: from 4,000 points to 100,000 This is the piece I'm most attached to, because the answer wasn't the one I started with — and the map only became fast once I accepted I'd found **two** bottlenecks, not one. ### What was actually broken The map had a hard ceiling at roughly **4,000 points**, and past it the tab died. Two separate causes, sitting on top of each other: **On the React side**, every point rendered its full detail component. Four thousand points meant roughly four thousand detail `div`s mounted at once. That filled memory, and then it crashed. **On the map side**, every point was drawn as its own marker. There was no clustering at all, and no separation of points by zoom scale — so a sudden zoom out, with the whole country in frame, still meant every single marker being drawn. Neither layer had any notion that the user can only actually look at a small part of this at a time. ### Why it became urgent For a long time this didn't hurt, because most customers weren't using the feature at scale. Then **Uber and Tata wanted to use it** — and the 4,000-point ceiling went from a known rough edge to a blocker in front of two of our largest accounts. It had to be fixed, quickly. That's when I picked it up. ### The wrong turn **First attempt: virtualization.** I started with the React side, because that was the crash. Only the components inside the visual window get rendered; as the user moves across the map, components outside the visible window are never created. It worked — for the crash. The page **stopped crashing** and got faster. But the map still felt wrong. **Zooming in and out was still janky.** Move the mouse and you could physically feel the drag. The page existed now, but it wasn't usable. ### The dig This is the part worth writing down. It would have been easy to conclude that virtualization was the answer and that what remained was just "the map being slow." Instead I went deeper — and found the two problems were sitting on top of each other. Virtualization was *a* bottleneck, not *the* bottleneck. Neither fix alone was going to make this fast, which is exactly why the first improvement felt like progress and still felt broken. The test that proved it: **when I stopped passing points to the map, the drag disappeared.** That located the second bottleneck precisely. It wasn't React any more — it was the sheer number of markers I was handing to the map itself. ### The fix: clustering by zoom level If the map can't take 100,000 markers, then the map should never be given 100,000 markers. It should be given whatever is meaningful **at the current zoom level** — because at most zoom levels, individual points aren't even legible. So I built a zoom-based cluster hierarchy: | Zoom level | What the map is handed | | --- | --- | | All the way out (world in frame) | A single roll-up — *India, 2,000 points* | | One level in | **State-level** clusters | | Further in | **District-level** clusters | | Close in | **Individual points per city** — how many, and exactly where | At every zoom level, the map is handed a small number of things to draw. The 100,000 points still exist in the data — they're just never all on screen at once, because they never needed to be. Alongside this came **layering**: separating individual points, clusters and sites into their own layers rather than throwing everything into one undifferentiated pile of markers. That separation is what makes the zoom rules above expressible at all — each layer can appear, disappear and update on its own terms. ### Debouncing the cluster computation Building clusters is real computation, and zooming is a continuous gesture. The naive version recalculates on **every frame** of a zoom — which just moves the cost from the map's rendering into my own JavaScript. So I debounced it: the cluster computation runs based on how much zooming has actually happened in a given interval, after the gesture settles rather than during every step of it. The user is mid-zoom; they don't need a recomputed cluster set at each intermediate level, they need one for where they land. ### Why clustering happened on the frontend The backend was sending all the data as-is, and the backend team was fully packed at that point — there was no room to change what was being sent. So the clustering was done entirely on the frontend. That was a constraint, not a preference, but it shaped the solution: everything above happens client-side, on data that arrives raw.
Outcome

Before: the page crashed outright past a few thousand points, and even when it survived, the drag was constant. After: it became a completely ordinary page — fast, and smooth to zoom exactly like any other map service. The ceiling went from 4,000 points to 100,000, and the experience at 100,000 is the one you'd expect from a map, not from a stress test.

> The lesson wasn't clustering or virtualization. It was that **the first fix > working is not evidence that the first fix was the whole answer.** > Virtualization stopped the crash and improved the page, and that success is > exactly what would have made most people stop. The map only got fast because I > kept going after the numbers had already improved. --- ## What I take away from this ### The through-line If I had to name the one thing I'm good at, it's **finding the actual problem**. Someone comes to me and says "this page isn't working." That sentence contains no information. My job starts there — working out *why* it isn't working, whether that's the performance of a page or the performance of a map. Virtualization, layering, clustering: those were the answers, but the value wasn't in knowing those techniques. It was in isolating the exact problem each one was the answer to — and, in the map's case, in not stopping when the first answer already looked like a win. The other half is that I can build a project from the base — given a clear requirement and a clear picture of how far it has to go — and build it to be scalable. On the frontend, scalable means something specific to me: - **performance is good** - **good practices are followed** - **state management is correct** - **components are structured in a correct, deliberate way** Those four aren't a checklist you apply at the end. They're the decisions that determine whether the thing survives its second year. ### What a year and a half taught me Most of what I learned came from working inside a **large codebase with a lot of dependencies**, where different teams work on different products and your change is never only your change. What I learned was a way of approaching work that arrives badly defined. A weak, loosely-specified problem lands on your desk, and the job is to give it structure: 1. **Do a thorough architecture analysis of the current code** — understand what exists before proposing to change it. 2. **Propose a solution on the basis of that analysis**, not on the basis of what you'd do on a blank page. 3. **Implement it in phases.** 4. **Test it, and find the bugs.** 5. **Fix them.** 6. **If something is wrong, turn around and make the right decision** — rather than defending the first one because it's yours. 7. **Debug the existing problems**, so the same class of problem stops recurring. Point six took the longest to internalise, and the map project is a monument to it. Being willing to say *this isn't it, go back further* is worth more than being right the first time. ### The technical ground Set out plainly, across these four projects: - **Virtualization** — rendering only what's in the visual window, and not constructing what isn't - **Zoom-based clustering and layering** on a map, with debounced recomputation - **Configuration-driven UI** — JSON describing position, data, appearance and representation - **Type-based component mapping** — turning a config entry into the right rendered component - **Compound / sub-component systems** — ``, ``-style APIs - **Lazy importing and code splitting** — at component, section and route level - **Web performance** — image optimization, bundle and chunk reduction, tree-shaking, CDN delivery, LCP/FCP work - **State management** at application scale - **Debugging** — the part that made all of the above possible ## Common questions ### How do you render 100,000 points on a map without the browser freezing? Never hand the map 100,000 markers. We took a map that crashed past roughly 4,000 points to a smooth 100,000 by building a zoom-based cluster hierarchy: fully zoomed out the map is handed a single country roll-up, then state-level clusters, then district-level, and only when you are close in does it get individual points. At every zoom level the map draws a small number of things, and the points that are not legible at that zoom are never sent. Clusters, individual points and sites also live in separate layers, which is what makes the per-zoom rules expressible at all. ### Why is a map still janky after virtualizing the list of points? Because there were two bottlenecks stacked on each other, and virtualizing only removed one. Windowing the React components stopped the page crashing and made it faster, but zooming still dragged. The test that located the real one was to stop passing points to the map at all — the drag vanished, which proved the remaining cost was the number of markers handed to the map itself, not React. Fixing one of two stacked bottlenecks feels like progress and still feels broken; that is the signal to keep digging rather than conclude the library is just slow. ### How do you cut a landing page load time from 2.5 seconds to under a second? Four things, and they compound rather than add. Get image sizes down and stop serving the wrong-sized image per breakpoint, which is what was causing the layout shift and the slow paint. Lazy-import the components that are not on screen at first load instead of bundling them into the initial payload. Put assets on a CDN, which matters as soon as your clients are in different regions from your origin. And tree-shake the unused CSS and JS, because a smaller file served from a nearer edge wins twice. Image and bundle work was the LCP and FCP work — they were not separate tasks. ---