{"id":32972,"date":"2023-10-13T21:12:00","date_gmt":"2023-10-13T21:12:00","guid":{"rendered":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint"},"modified":"2025-05-09T10:18:58","modified_gmt":"2025-05-09T10:18:58","slug":"llmops-platform-blueprint","status":"publish","type":"post","link":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint","title":{"rendered":"LLMOps blueprint for closed-source large language models"},"content":{"rendered":"<p><b style=\"display:none!important;\" itemprop=\"image\" itemscope itemtype=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/schema.org\/ImageObject\"><\/b><\/p>\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-11jzjcui\"><p class=\"uagb-heading-text\">Building solutions using closed-source large language models (LLMs), including models like GPT-4 from OpenAI, or PaLM2 from Google, is a markedly different process to creating private machine learning (ML) models, so traditional <a href=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/solutions\/mlops\">MLOps<\/a> playbooks and best practices might appear irrelevant when applied to LLM-centric projects. And indeed, many companies currently approach LLM projects as greenfield developments that cannot fully reuse the existing MLOps infrastructure, and that do not have a clear strategy for creating any new reusable infrastructure. This approach invariably results in the accumulation of significant technical debt for new LLM-based solutions, elevated infrastructure costs, operational inefficiencies, and security and safety breaches.\u00a0\u00a0\u00a0<\/p><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-ridp8k5w\"><p class=\"uagb-heading-text\">In this article, we describe a reusable infrastructure (platform) for solutions that are powered by closed-source LLMs. This platform aims to streamline the development and operation of LLM-based solutions by standardizing services, tools, and other generic components across the applications. We focus specifically on closed-source LLMs accessible via APIs (such as OpenAI GPT-4, Google PaLM2, or Anthropic Claude). It\u2019s worth noting that a platform for training and hosting open-source-based LLMs is a different topic that requires a separate discussion.<\/p><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-i8wfsxod\" id=\"llmops-capabilities-and-features\"><h2 id='llmops-capabilities-and-features'  id=\"boomdevs_1\" class=\"uagb-heading-text\">LLMOps capabilities and features<\/h2><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-bwg81spj\"><p class=\"uagb-heading-text\">For the sake of specificity, let us assume a retrieval-augmented generation (RAG) application that enables internal or external users to query a collection of documents (knowledge base) via a conversational interface. Let us also assume that this application has access to <em>tools<\/em>, including systems such as relational databases or search engines that can be queried or updated using API calls or SQL queries. A general layout of such a <em>RAG system with tools<\/em> is shown in the figure below:<\/p><\/div>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1730\" height=\"340\" src=\"data:image\/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==\" data-lazy-type=\"image\" data-aload=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-1.jpg\" alt=\"General layout of a retrieval-augmented generation (RAG) application with tools\" class=\"lazy wp-image-32969\" data-aload-srcset=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-1.jpg 1730w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-1-300x59.jpg 300w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-1-1024x201.jpg 1024w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-1-768x151.jpg 768w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-1-1536x302.jpg 1536w\" data-aload-sizes=\"auto, (max-width: 1730px) 100vw, 1730px\" \/><noscript><img loading=\"lazy\" decoding=\"async\" width=\"1730\" height=\"340\" src=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-1.jpg\" alt=\"General layout of a retrieval-augmented generation (RAG) application with tools\" class=\"wp-image-32969\" srcset=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-1.jpg 1730w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-1-300x59.jpg 300w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-1-1024x201.jpg 1024w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-1-768x151.jpg 768w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-1-1536x302.jpg 1536w\" sizes=\"auto, (max-width: 1730px) 100vw, 1730px\" \/><\/noscript><\/figure>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-xv6cg0u3\"><p class=\"uagb-heading-text\">This architecture can support a wide range of business scenarios. Consider the following examples:<\/p><\/div>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\nA vendor of a software platform for financial advisors needs to build a knowledge management system that allows the advisors to ask questions about the platform, financial products, and tax policies. Under the hood, the system searches through an extensive collection of documents and other unstructured data.&nbsp;\n<\/li>\n\n\n\n<li>\nA wholesaler of industrial goods needs to build a knowledge management system that allows customers to ask questions about products and browse product details. Internally, the system queries product specifications submitted by the supplier in PDF format.&nbsp;\n<\/li>\n\n\n\n<li>\nA banking company needs to provide its associates with a user-friendly business intelligence (BI) solution for querying stock data using natural language questions like \u201cWhat were the daily trading volumes for Alphabet stock over the last two weeks?\u201d. Internally, the application uses an LLM to generate SQL and query a relational database (which we view as a tool).&nbsp;\n<\/li>\n<\/ul>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-swt7xpws\"><p class=\"uagb-heading-text\">Although the architecture outlined above is relatively straightforward, it necessitates the resolution of recurrent problems that manifest in virtually all RAG applications, as well as in LLM applications in general:<\/p><\/div>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\n<strong>Document storing and searching:<\/strong> RAG architecture requires a <em>vector database<\/em> for indexing knowledge base documents and their components, while enabling nearest neighbor search. This capability is commonly implemented using embedding-based search which, in turn, requires a model or service for computing text embeddings.\n<\/li>\n\n\n\n<li>\n<strong>Document preprocessing: <\/strong>The documents should typically be preprocessed to extract text from structured formats like PDF or mask-sensitive information such as PII data.&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Document indexing:<\/strong> The preprocessed documents are indexed in the vector database, which can be a relatively complex process that requires tracking the removed\/updated\/deleted documents in the knowledge base, enriching documents with metadata such as document type and access groups, splitting documents into chunks, and computing embeddings.\n<\/li>\n\n\n\n<li>\n<strong>LLM gateway:<\/strong> The platform should support integrations with various LLM providers and models, including text generation and embedding computing models. The interaction gateway can also provide features like rate limiting, autoretry, etc.&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Model management:<\/strong> Providers of closed-source LLMs commonly provide the ability to fine-tune models on proprietary data, necessitating proper management of the corresponding datasets and collections of fine-tuned models.\n<\/li>\n\n\n\n<li>\n<strong>Prompt management:<\/strong> LLM-based applications tend to use a large number of prompts that directly or indirectly control system behavior and user experience. The applications should use a unified framework for managing and versioning prompts.\n<\/li>\n\n\n\n<li>\n<strong>Observability: <\/strong>The production instance of the application needs to be continuously monitored to detect technical and functional issues in early stages. In particular, the problem of text generation quality should be continuously evaluated.&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Cost efficiency:<\/strong> The platform should provide caching, cost tracking, and other services that optimize the cost of LLM usage.&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Performance and latency:<\/strong> Text generation using LLMs is not a particularly fast operation, necessitating means for improving latency and other performance properties of the solution.&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Safety and compliance:<\/strong> The platform should provide a mechanism for detecting and preventing LLM hallucinations, toxic responses, and other threats to a safe and consistent user experience. The platform should also provide protection for tools\/systems that are controlled by LLMs.&nbsp;&nbsp;&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Protection against bad actors:<\/strong> Security and data privacy are very important aspects of LLM-based solutions, and the platform should provide protection on both the LLM and user sides.\n<\/li>\n<\/ul>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-rm3i1xqe\"><p class=\"uagb-heading-text\">In the next sections, we will demonstrate how to develop an LLMOps platform that addresses these concerns. We start with a minimal solution that is required to support basic RAG applications and add more advanced capabilities step by step.<\/p><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-rty8ar4m\" id=\"platform-architecture\"><h2 id='platform-architecture'  id=\"boomdevs_2\" class=\"uagb-heading-text\">Platform architecture<\/h2><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-p936a8tp\"><p class=\"uagb-heading-text\">A basic RAG application usually includes two flows: indexing and querying. The indexing flow can be viewed as a data (document) processing pipeline that performs the following steps:<\/p><\/div>\n\n\n\n<ol class=\"wp-block-list\">\n<li>\nDocument preprocessing such as PII data masking.\n<\/li>\n\n\n\n<li>\nSplitting the documents into manageable chunks that fit the LLM context, and attributing metadata to each chunk.\n<\/li>\n\n\n\n<li>\nComputing embeddings for the chunks.\n<\/li>\n\n\n\n<li>\nSaving the chunks indexed by embeddings to the vector database.\n<\/li>\n<\/ol>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-p2vtwh10\"><p class=\"uagb-heading-text\">The querying flow is executed in run time for each request (e.g. each turn in a chatbot dialog) and includes the following steps:<\/p><\/div>\n\n\n\n<ol class=\"wp-block-list\">\n<li>\nThe application creates a query context (e.g. by concatenating all messages in the dialog history including the latest input from the user) and converts it into a standalone question using the text-generating LLM.&nbsp;\n<\/li>\n\n\n\n<li>\nThe embedding for this standalone question is computed using an embedding model.\n<\/li>\n\n\n\n<li>\nRelevant document chunks are searched in the vector database based on the question embedding (nearest neighbor search).\n<\/li>\n\n\n\n<li>\nThe relevant chunks can be additionally filtered based on the metadata.\n<\/li>\n\n\n\n<li>\nThe chunks are retrieved, combined, and injected as the context, enriching the LLM\u2019s ability to produce a relevant and well-grounded response.\n<\/li>\n<\/ol>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-sz37eo2n\"><p class=\"uagb-heading-text\">The minimal set of components required for these two flows includes an LLM API for text generation, an LLM API for embedding computing, a data (document) indexing pipeline, and a vector database. These essential components are shown in green in the platform architecture diagram presented below.<\/p><\/div>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1730\" height=\"1106\" src=\"data:image\/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==\" data-lazy-type=\"image\" data-aload=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-2.jpg\" alt=\"LLMops platform architecture showing LLM API for text generation, LLM API for embedding computing, a data (document) indexing pipeline, and a vector database in green\" class=\"lazy wp-image-32970\" data-aload-srcset=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-2.jpg 1730w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-2-300x192.jpg 300w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-2-1024x655.jpg 1024w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-2-768x491.jpg 768w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-2-1536x982.jpg 1536w\" data-aload-sizes=\"auto, (max-width: 1730px) 100vw, 1730px\" \/><noscript><img loading=\"lazy\" decoding=\"async\" width=\"1730\" height=\"1106\" src=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-2.jpg\" alt=\"LLMops platform architecture showing LLM API for text generation, LLM API for embedding computing, a data (document) indexing pipeline, and a vector database in green\" class=\"wp-image-32970\" srcset=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-2.jpg 1730w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-2-300x192.jpg 300w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-2-1024x655.jpg 1024w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-2-768x491.jpg 768w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-2-1536x982.jpg 1536w\" sizes=\"auto, (max-width: 1730px) 100vw, 1730px\" \/><\/noscript><\/figure>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-y89ummx2\"><p class=\"uagb-heading-text\">In this architecture, we assume that interoperability with multiple LLM providers is ensured by the LLM framework used by the application (e.g. LangChain). The LLM framework can also provide built-in document preprocessing components, integrations with vector databases, and large building blocks that implement complete RAG and Agent flows.<\/p><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-fybe0s3o\"><p class=\"uagb-heading-text\">However, the green components alone do not address all the challenges and concerns that we outlined in the previous section. For this reason, LLMOps platforms usually include additional services that provide more advanced functionality. These services are depicted in yellow and blue in the above diagram, and we discuss them one by one in the next sections. We also discuss how these components can be integrated with the basic flow (e.g. how document preprocessing can be coordinated with query-time post-processing). <\/p><\/div>\n\n\n<p><b><strong style=\"white-space: pre-wrap;\">Example technology stack<\/strong><\/b><\/p>\n<p><b><strong style=\"white-space: pre-wrap;\">&#8211; Data indexing:<\/strong><\/b><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Spark, LangChain, Airflow<\/em><\/i><br \/><b><strong style=\"white-space: pre-wrap;\">&#8211; Vector database:<\/strong><\/b><i><em class=\"italic\" style=\"white-space: pre-wrap;\">ChromaDB, Qdrant, Milvus, Pinecone, or FAISS<\/em><\/i><br \/><b><strong style=\"white-space: pre-wrap;\">&#8211; Application stack:<\/strong><\/b><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Python, LangChain, LlamaIndex, FastAPI<\/em><\/i><\/p>\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-1e7xyoiq\" id=\"caching-and-streaming\"><h2 id='caching-and-streaming'  id=\"boomdevs_3\" class=\"uagb-heading-text\">Caching and streaming<\/h2><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-7w0o2l2s\"><p class=\"uagb-heading-text\">Text generation using LLMs is a relatively slow operation which often becomes a major issue from the user experience standpoint. This problem can be addressed using several techniques:<\/p><\/div>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\n<strong>Caching:<\/strong> The LLM responses can be cached to handle frequent queries\/contexts more efficiently. The most basic solution is to use standard caching components like Redis that perform an exact match between a new query and a cached query. However, this approach might be inefficient in LLM applications because of the high variability of queries. A more sophisticated solution lies in <em>semantic caching,<\/em> which employs embedding-based search to find cached queries that are similar to the new one. Internally, a semantic cache can rely on a vector database.&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Streaming:<\/strong> User experience can be improved by printing out the generated text token by token (similar to the approach employed in the ChatGPT web interface) instead of waiting for a complete response in a blocking mode. This technique is commonly referred to as streaming. Streaming APIs are offered by most LLM providers, but creating a streaming application requires propagating near real-time token streams through all layers including backend services and frontend UI. Consequently, streaming needs to be incorporated in the solution design and appropriate orchestration and UI frameworks should be selected.\n<\/li>\n<\/ul>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-j7czvdx8\"><p class=\"uagb-heading-text\">The caching layer can also help to reduce the cost of the LLM services, although the efficiency of this solution can greatly vary depending on the application.<\/p><\/div>\n\n\n<p><b><strong style=\"white-space: pre-wrap;\">Example technology stack<\/strong><\/b><\/p>\n<p><b><strong style=\"white-space: pre-wrap;\">&#8211; Exact match caching:<\/strong><\/b><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Redis<\/em><\/i><br \/><b><strong style=\"white-space: pre-wrap;\">&#8211; Semantic caching:<\/strong><\/b><i><em class=\"italic\" style=\"white-space: pre-wrap;\">GPTCache + vector database<\/em><\/i><br \/><b><strong style=\"white-space: pre-wrap;\">&#8211; Streaming:<\/strong><\/b><i><em class=\"italic\" style=\"white-space: pre-wrap;\">OpenAI Streaming, LangChain (supports streaming), Streamlit Chat (supports streaming)<\/em><\/i><\/p>\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-uusnm7ol\" id=\"feature-store\"><h2 id='feature-store'  id=\"boomdevs_4\" class=\"uagb-heading-text\">Feature store<\/h2><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-b1g3vqz2\"><p class=\"uagb-heading-text\">LLM-powered applications often need to look up customer data and other contextual information. For example, a customer assistant app might fetch information about the loyalty tier and engagement level from a customer profile to personalize generated messages.<\/p><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-o2jmkjyg\"><p class=\"uagb-heading-text\">In general, LLM-powered applications can be integrated with any source of contextual data such as customer data platforms (CDPs). However, it is often beneficial to use a feature store, a concept that is well-known in traditional MLOps, for managing and sourcing such data. The rationale behind this choice is manifold:<\/p><\/div>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\n<strong>Tailored for low-latency transactions:<\/strong> Feature stores are purpose-built to provide low-latency transactional access to large collections of features including basic attributes and derived values such as propensity scores. This design effectively mitigates scalability bottlenecks, ensuring swift and efficient data retrieval.\n<\/li>\n\n\n\n<li>\n<strong>Interoperability with ML models:<\/strong> Feature stores provide interoperability between in-house ML models (e.g. personalization or recommendation models) and LLM applications. If the company already uses a feature store, it is easy for LLM applications to tap into this infrastructure.\n<\/li>\n\n\n\n<li>\n<strong>Continuous feature updates:<\/strong> Feature stores are typically connected with supplementary infrastructure elements like data collection pipelines and scoring models, ensuring that the features remain consistently updated.\n<\/li>\n<\/ul>\n\n\n<p><b><strong style=\"white-space: pre-wrap;\">Example technology stack<\/strong><\/b><\/p>\n<p><i><em class=\"italic\" style=\"white-space: pre-wrap;\">&#8211; Feast, Vertex AI Feature Store, or Databricks Feature Store<\/em><\/i><\/p>\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-2z1wssov\" id=\"prompt-management\"><h2 id='prompt-management'  id=\"boomdevs_5\" class=\"uagb-heading-text\">Prompt management<\/h2><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-viktvx34\"><p class=\"uagb-heading-text\">Prompts are among the most important building blocks for <a href=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/driving-business-success-with-generative-ai-techniques-for-value-driven-transformation\">Generative AI (GenAI) applications<\/a>, second only to LLMs themselves. Why? Because application behavior and the majority of the business logic are encoded in the prompts. Complex applications can use tens or hundreds of different prompts, many of which are parametrized with dynamic data or even generated by LLMs.<\/p><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-qcs3jl1g\"><p class=\"uagb-heading-text\">In small standalone applications, prompts can be hard coded or stored in configuration files. However, complex enterprise applications might require more solid infrastructure for prompt management that not only enables the dynamic alteration of application behavior, but also expedites the resolution of user experience issues, facilitates AB testing, allows for test changes and fixes before production deployment, and even automates prompt testing on multiple models and sets of input parameters. These functions can be consolidated in a separate prompt management system. Prompt management services are available in some ML platforms and as standalone products.<\/p><\/div>\n\n\n<p><b><strong style=\"white-space: pre-wrap;\">Example technology stack<\/strong><\/b><\/p>\n<p><i><em class=\"italic\" style=\"white-space: pre-wrap;\">&#8211; PromptHub<\/em><\/i><\/p>\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-4hbwxg7e\" id=\"guardrails-safety-compliance-and-user-experience\"><h2 id='guardrails-safety-compliance-and-user-experience'  id=\"boomdevs_6\" class=\"uagb-heading-text\">Guardrails: Safety, compliance, and user experience<\/h2><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-f7qeu4g4\"><p class=\"uagb-heading-text\">The pursuit of user experience quality and safety within LLM-based applications presents a formidable challenge, owing to a multitude of factors, including the following:<\/p><\/div>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\n<strong>Inherent complexity of generated text evaluation: <\/strong>Assessing the quality, usefulness, compliance, and safety of generated text is inherently difficult to evaluate in general.\n<\/li>\n\n\n\n<li>\n<strong>Versatility of conversational systems:<\/strong> Conversational systems are extremely versatile and flexible, making it difficult or impossible to perform comprehensive testing.\n<\/li>\n\n\n\n<li>\n<strong>Integration with vector search:<\/strong> Vector search used in RAG adds an additional layer of complexity and uncertainty.\n<\/li>\n\n\n\n<li>\n<strong>Continuous LLM updates:<\/strong> LLM providers continuously update their products, causing shifts in application behavior that can prove unpredictable.\n<\/li>\n\n\n\n<li>\n<strong>Testability of complex LLM chains:<\/strong> Given their intricate interdependencies, complex LLM chains are difficult to test and debug.\n<\/li>\n<\/ul>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-1zmu5zrd\"><p class=\"uagb-heading-text\">These challenges can be addressed at different levels of the LLM stack including datasets for LLM training, fine-tuning, and various interceptors on the LLM provider and application sides. From the LLMOps perspective, request\/response interceptors on the application side are a very important and powerful technique. These interceptors, commonly called <em>guardrails<\/em>, can perform a broad range of checks and make corrections to steer the application behavior, improve user experience, and prevent safety and compliance issues. Examples of safety and compliance guardrails include the following:<\/p><\/div>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\n<strong>Toxicity assessment:<\/strong> This is a component that evaluates text for the presence of violence, harmful, or offensive content, generating scores (e.g. hate_speech = 0.28) that trigger mitigating actions. The toxicity checks can be performed for both user input and LLM output.\n<\/li>\n\n\n\n<li>\n<strong>Topic bans:<\/strong> This is a component that detects certain topics and takes mitigating actions. For example, it can detect politics-related topics in the user input (e.g. \u201cWhat do you think of the president?\u201d) and return a blocking response (e.g. \u201cI&#8217;m a shopping assistant, I don&#8217;t like to talk about politics\u201d). This check can be applied to both user input and LLM output.&nbsp;&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Relevance validation:<\/strong> This is a guardrail that validates the generated response as relevant to the input question and context, often measured through similarity scores between the prompt and response. For example, the score will be low for the question \u201cWhat is the current sales tax in California?\u201d when the answer is \u201cYangtze River is the longest river in both China and Asia\u201d.&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Contradiction assessment:<\/strong> This is a component that verifies the absence of LLM output self-contradictions, contradictions to the input, or contradictions with established facts.&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Hallucinations and factuality detection:<\/strong> This guardrail detects hallucinations and non-factual statements in the LLM output. Detecting hallucinations is generally a challenging task that can be approached in many different ways. One practical approach for closed-source LLMs is to generate multiple responses for the same query and validate their agreement.\n<\/li>\n<\/ul>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-ruxunzvx\"><p class=\"uagb-heading-text\">The standard checks described above require limited or no configuration and can be implemented using off-the-shelf libraries. However, you might also need to create more advanced guardrails that combine custom checks and mitigating actions into complex flows. Such custom guardrails might not even be directly related to safety or compliance, but steer other aspects of application behavior that require more flexible frameworks.<\/p><\/div>\n\n\n<p><b><strong style=\"white-space: pre-wrap;\">Example technology stack<\/strong><\/b><\/p>\n<p><b><strong style=\"white-space: pre-wrap;\">&#8211; Standard guardrails:<\/strong><\/b><i><em class=\"italic\" style=\"white-space: pre-wrap;\">LLM-Guard, OpenAI Moderation APIs<\/em><\/i><br \/><b><strong style=\"white-space: pre-wrap;\">&#8211; Custom flows:<\/strong><\/b><i><em class=\"italic\" style=\"white-space: pre-wrap;\">NVIDIA NeMo, custom guards using Hugging Face model<\/em><\/i><\/p>\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-bsg4uhvk\" id=\"guardrails-tools\"><h2 id='guardrails-tools'  id=\"boomdevs_7\" class=\"uagb-heading-text\">Guardrails: Tools<\/h2><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-goizc7o2\"><p class=\"uagb-heading-text\">Applications that integrate with external tools invariably demand specialized guardrails. For example, an application that queries a relational database using text-to-SQL generation needs to validate that the generated SQL is valid, execute it, analyze the response returned by the database, fix the SQL query in case of errors, and repeat until the required data are fetched or retry limits are reached. The guardrails can become even more critical and sophisticated when the application can update data or change the tool state.<\/p><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-h913d5m1\" id=\"guardrails-security-and-privacy\"><h2 id='guardrails-security-and-privacy'  id=\"boomdevs_8\" class=\"uagb-heading-text\">Guardrails: Security and privacy<\/h2><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-m3tr4oww\"><p class=\"uagb-heading-text\">While the previous section delved into the multifaceted challenges of quality and safety, the realm of LLM usage introduces a secondary concern of equal significance: security and privacy. Similar to the management of quality and safety, security and privacy concerns demand an equally comprehensive approach, spanning various layers of the LLM framework. This includes establishing contractual terms with LLM providers, implementing training and RAG data preprocessing and minimization methods, and orchestrating real-time guardrails that scan user inputs and generated outputs.&nbsp;<\/p><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-ns8e4xlg\"><p class=\"uagb-heading-text\">One possible way of planning the security strategy is to list different classes of assets (data and systems) that we need to protect and possible threats, and make sure that each asset-threat combination is covered using one or several methods.&nbsp;<\/p><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-n4z1lakl\"><p class=\"uagb-heading-text\">An example of such planning is presented in the table below. It is important to note that this is a simplified illustration that doesn&#8217;t account for open-source or private LLMs and associated techniques such as differential privacy.<\/p><\/div>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1730\" height=\"820\" src=\"data:image\/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==\" data-lazy-type=\"image\" data-aload=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-3.jpg\" alt=\"Table showing a security strategy with a list of different classes of assets (data and systems) and possible threats. Each asset-threat combination is covered using one or several methods\" class=\"lazy wp-image-32971\" data-aload-srcset=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-3.jpg 1730w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-3-300x142.jpg 300w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-3-1024x485.jpg 1024w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-3-768x364.jpg 768w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-3-1536x728.jpg 1536w\" data-aload-sizes=\"auto, (max-width: 1730px) 100vw, 1730px\" \/><noscript><img loading=\"lazy\" decoding=\"async\" width=\"1730\" height=\"820\" src=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-3.jpg\" alt=\"Table showing a security strategy with a list of different classes of assets (data and systems) and possible threats. Each asset-threat combination is covered using one or several methods\" class=\"wp-image-32971\" srcset=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-3.jpg 1730w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-3-300x142.jpg 300w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-3-1024x485.jpg 1024w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-3-768x364.jpg 768w, https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/01\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-img-3-1536x728.jpg 1536w\" sizes=\"auto, (max-width: 1730px) 100vw, 1730px\" \/><\/noscript><\/figure>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-kcwv2ypz\"><p class=\"uagb-heading-text\">Similar to safety and compliance, many security and privacy threats can be addressed using real-time guardrails. Standard examples of security guardrails include the following:<\/p><\/div>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\n<strong>Anonymization\/deanonymization:<\/strong> This guardrail consists of two parts. The first one is the input interceptor that detects sensitive data elements such as names, addresses, and other PII data, and replaces them with surrogate tokens. The second is the output interceptor that performs the reverse mapping. In some cases, the anonymization operation can be made irreversible.&nbsp;&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Prompt injection:<\/strong> This is a component that detects malicious prompts that attempt to override system prompts, hijack the conversation context, or instruct the LLM to perform unintended actions (jailbreak). This guardrail is particularly important for LLM agents that control internal or external systems via API.\n<\/li>\n\n\n\n<li>\n<strong>Secret detection:<\/strong> This is a scanner that detects login, password, credit card numbers and other sensitive data. This check can be applied to both user input and LLM output.&nbsp;&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>URL detection:<\/strong> This scanner detects malicious URLs in user input and LLM output.\n<\/li>\n\n\n\n<li>\n<strong>Rate limiting:<\/strong> This guardrail throttles input querying and LLM invocation rates with the goal of preventing outages, DoS attacks, and excessive (and costly) LLM calls due to defects.\n<\/li>\n<\/ul>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-5luo8wrg\"><p class=\"uagb-heading-text\">One of the key considerations for designing or selecting a security guardrail framework is the ability to efficiently react to new threats and breaches. From that perspective, security guardrails should follow the protocols and best practices used in the cybersecurity industry.<\/p><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-0b0knswm\"><p class=\"uagb-heading-text\">It is worth noting that some of the above guardrails require coordination between the document preprocessing pipeline and query-time processing. For example, a RAG system can prevent the LLM provider from seeing the PII data using the following approach:<\/p><\/div>\n\n\n\n<ol class=\"wp-block-list\">\n<li>\nThe PII data can be replaced with surrogate tokens as a part of the preprocessing stage.\n<\/li>\n\n\n\n<li>\nThe mapping between actual PII values and tokens is saved in a database.\n<\/li>\n\n\n\n<li>\nReverse mapping is performed in query time to produce the final response for the user. \n<\/li>\n<\/ol>\n\n\n<p><b><strong style=\"white-space: pre-wrap;\">Example technology stack<\/strong><\/b><\/p>\n<p><i><em class=\"italic\" style=\"white-space: pre-wrap;\">&#8211; LLM-Guard, Guardrail ML, custom guards using Hugging Face model<\/em><\/i><\/p>\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-7rdghmrt\" id=\"observability\"><h2 id='observability'  id=\"boomdevs_9\" class=\"uagb-heading-text\">Observability<\/h2><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-pngzksr4\"><p class=\"uagb-heading-text\">The safety, compliance, and security challenges discussed in the previous sections underscore the importance of observability, which is another well-known concept from traditional MLOps. In the case of LLMOps, observability use cases include the following:<\/p><\/div>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\n<strong>Prompt analytics:<\/strong> The prompts received from users or external systems should be logged and analyzed. In particular, the prompts can be clustered and visualized to better understand the application usage.\n<\/li>\n\n\n\n<li>\n<strong>Problematic prompts:<\/strong> Prompts that deviate from the typical clusters can be flagged as outliers.\n<\/li>\n\n\n\n<li>\n<strong>User feedback capturing:<\/strong> As we discussed earlier, the quality of generated text can be challenging to evaluate. The quality analysis and detection of problematic prompts can be facilitated by capturing implicit or explicit user feedback (e.g. thumbs up\/down).&nbsp;&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Automatic checks:<\/strong> Continuous LLM updates can be countered with automatic quality, compliance, and safety checks.&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Monitoring: <\/strong>LLM-backed applications require tracking metrics such as cache hit ratio and throughputs\/latencies at different stages of the LLM chains.\n<\/li>\n\n\n\n<li>\n<strong>Logging:<\/strong> Logging requests, responses, prompts, and automatic checks enable auditability, and support optimization and troubleshooting.\n<\/li>\n<\/ul>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-slhrg9n7\"><p class=\"uagb-heading-text\">These and other observability capabilities can be implemented using specialized products or custom tools. Observability components should also be integrated with data lakes\/warehouses and general-purpose analytics and BI tools to support deep analysis and application optimization.<\/p><\/div>\n\n\n<p><b><strong style=\"white-space: pre-wrap;\">Example technology stack<\/strong><\/b><\/p>\n<p><i><em class=\"italic\" style=\"white-space: pre-wrap;\">&#8211; Arize<\/em><\/i><\/p>\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-w6ekyoqf\" id=\"model-management\"><h2 id='model-management'  id=\"boomdevs_10\" class=\"uagb-heading-text\">Model management<\/h2><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-0xom3626\"><p class=\"uagb-heading-text\">In the previous sections, we implicitly assumed immutable foundation models. However, most LLM vendors provide the ability to fine-tune the foundation models via specialized APIs. This creates an additional layer of complexity from the LLMOps perspective. More specifically, we need to provide the following capabilities:<\/p><\/div>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\n<strong>Training dataset management:<\/strong> Fine-tuning hinges on the availability of training datasets, which demand data cataloging and tracking, mirroring the conventional practices applied to datasets in traditional MLOps.&nbsp;\n<\/li>\n\n\n\n<li>\n<strong>Model registry:<\/strong> As the number of fine-tuned models grows, maintaining a registry of such models with the associated metadata, such as training parameters, becomes an important concern. This is also a well-understood problem in traditional MLOps.&nbsp;\n<\/li>\n<\/ul>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-e53tdz4y\"><p class=\"uagb-heading-text\">Conceptually, these capabilities can be implemented using the same tools as in traditional MLOps, but LLMOps imposes several unique challenges. For example, fine-tuned models can be hosted across multiple LLM providers making it more difficult to create a unified registry. Another example is model metrics\u2013quality metrics are an integral part of the metadata in traditional MLOps, but automatic evaluation of the quality metrics in LLMOps is much more challenging, although possible.<\/p><\/div>\n\n\n<p><b><strong style=\"white-space: pre-wrap;\">Example technology stack<\/strong><\/b><\/p>\n<p><b><strong style=\"white-space: pre-wrap;\">&#8211; Training dataset management:<\/strong><\/b><i><em class=\"italic\" style=\"white-space: pre-wrap;\">DataHub<\/em><\/i><br \/><b><strong style=\"white-space: pre-wrap;\">&#8211; Model repository:<\/strong><\/b><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Vertex AI, MLflow<\/em><\/i><\/p>\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-6p2v6wa9\" id=\"discussion\"><h2 id='discussion'  id=\"boomdevs_11\" class=\"uagb-heading-text\">Discussion<\/h2><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-en71xkzx\"><p class=\"uagb-heading-text\">Throughout the previous sections, we have described the LLMOps platform specifically for closed-source LLMs, but is it reasonable to build applications using only closed-source models? From this strategic perspective, it is generally recommended to consider individual functions that LLMs perform, and make design decisions for each function separately:<\/p><\/div>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\nAuxiliary functions such as embedding computing, PII data detection and masking, and some other operations that are performed by guardrails do not necessarily require cutting-edge LLMs and can be implemented using small models. Consequently, these functions are good candidates for implementing open-source, privately deployed, and fine-tuned models.\n<\/li>\n\n\n\n<li>\nText generation and reasoning capabilities typically require state-of-the-art LLMs which makes them more difficult to implement using the open-source approach. Using privately deployed, closed-source LLMs or open-source LLMs for text generation is a major strategic decision that is informed by functional, security, budgeting, operational, and business considerations.\n<\/li>\n<\/ul>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-sslwhon7\"><p class=\"uagb-heading-text\">For these reasons, combining closed-source and open-source models can be particularly advantageous. For example, security concerns can be effectively addressed by using privately deployed open-source models for detecting and masking sensitive data, while subsequently, the actual business flow is executed using a public closed-source LLM based on the masked data.&nbsp;<\/p><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-zmkb3n34\" id=\"conclusion\"><h2 id='conclusion'  id=\"boomdevs_12\" class=\"uagb-heading-text\">Conclusion<\/h2><\/div>\n\n\n\n<div class=\"wp-block-uagb-advanced-heading uagb-block-oddnv63m gd-links-underline-container\"><p class=\"uagb-heading-text\">GenAI adoption is expected to grow at a near-exponential pace in the next few years, and as the scale and complexity of GenAI solutions increase, the role of LLMOps will become increasingly important. Moreover, many pilot solutions that were developed in the early stages of GenAI adoption will ultimately require migration to solid, manageable, and cost-efficient LLMOps platforms. In this article, we outlined the functional and technical designs of such a platform specifically for closed-source LLMs, and provided mappings to well-established and emerging frameworks, libraries, and products that can be used for implementation.  If you&#8217;re looking to learn more about how we can assist you, be sure to check out our <strong><a href=\"\/https\/www.griddynamics.com\/services\/artificial-intelligence\">AI for Enterprise<\/a><\/strong> service page.<\/p><\/div>\n\n\n\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Building solutions using closed-source large language models (LLMs), including models like GPT-4 from OpenAI, or PaLM2 from Google, is a markedly different process to creating private machine learning (ML) models, so traditional MLOps playbooks and best practices might appear irrelevant when applied to LLM-centric projects. And indeed, many companies currently approach LLM projects as greenfield<\/p>\n","protected":false},"author":103,"featured_media":63225,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_eb_attr":"","_uag_custom_page_level_css":"","site-sidebar-layout":"no-sidebar","site-content-layout":"plain-container","ast-site-content-layout":null,"site-content-style":null,"site-sidebar-style":null,"ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"disabled","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":""},"categories":[225,222,223,260,230,243],"tags":[],"starter_kit_tag":[],"class_list":["post-32972","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-data-and-ml-platforms","category-articles","category-ai","category-cross-industry","category-generative-ai","category-llmops"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v26.8 (Yoast SEO v26.8) - https:\/\/round-lake.dustinice.workers.dev:443\/https\/yoast.com\/product\/yoast-seo-premium-wordpress\/ -->\n<title>LLMOps platform infrastructure for GenAI solutions<\/title>\n<meta name=\"description\" content=\"Learn how to migrate GenAI pilot projects to a robust LLMOps platform. Explore technical foundations, guardrails &amp; implementation insights for closed-source LLMs\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"LLMOps platform infrastructure for GenAI solutions\" \/>\n<meta property=\"og:description\" content=\"Discover the essential role of LLMOps in scaling GenAI solutions and how to migrate pilot projects to a robust LLMOps platform. Explore technical foundations, guardrails, and implementation insights for closed-source LLMs.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint\" \/>\n<meta property=\"og:site_name\" content=\"Grid Dynamics\" \/>\n<meta property=\"article:published_time\" content=\"2023-10-13T21:12:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2025-05-09T10:18:58+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/05\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-Cover-v1-1.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1300\" \/>\n\t<meta property=\"og:image:height\" content=\"460\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Dmitry Mezhensky\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:title\" content=\"LLMOps blueprint for closed-source large language models\" \/>\n<meta name=\"twitter:description\" content=\"Grid Dynamics - LLMOps blueprint for closed-source large language models - AI and data platforms, Articles, Artificial intelligence, Cross-industry, Generative AI, LLMOps - October 13, 2023\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Dmitry Mezhensky\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"17 minutes\" \/>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"LLMOps platform infrastructure for GenAI solutions","description":"Learn how to migrate GenAI pilot projects to a robust LLMOps platform. Explore technical foundations, guardrails & implementation insights for closed-source LLMs","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint","og_locale":"en_US","og_type":"article","og_title":"LLMOps platform infrastructure for GenAI solutions","og_description":"Discover the essential role of LLMOps in scaling GenAI solutions and how to migrate pilot projects to a robust LLMOps platform. Explore technical foundations, guardrails, and implementation insights for closed-source LLMs.","og_url":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint","og_site_name":"Grid Dynamics","article_published_time":"2023-10-13T21:12:00+00:00","article_modified_time":"2025-05-09T10:18:58+00:00","og_image":[{"width":1300,"height":460,"url":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2024\/05\/LLMOps-Blueprint-for-Closed-Source-Large-Language-Models-Cover-v1-1.jpg","type":"image\/jpeg"}],"author":"Dmitry Mezhensky","twitter_card":"summary_large_image","twitter_title":"LLMOps blueprint for closed-source large language models","twitter_description":"Grid Dynamics - LLMOps blueprint for closed-source large language models - AI and data platforms, Articles, Artificial intelligence, Cross-industry, Generative AI, LLMOps - October 13, 2023","twitter_misc":{"Written by":"Dmitry Mezhensky","Est. reading time":"17 minutes"},"schema":{"@context":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint#article","isPartOf":{"@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint"},"author":{"name":"Dmitry Mezhensky","@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/#\/schema\/person\/e12e66db4925524135888aea2e3d12eb"},"headline":"LLMOps blueprint for closed-source large language models","datePublished":"2023-10-13T21:12:00+00:00","dateModified":"2025-05-09T10:18:58+00:00","mainEntityOfPage":{"@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint"},"wordCount":3729,"publisher":{"@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/#organization"},"image":{"@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint#primaryimage"},"thumbnailUrl":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/05\/llmops-blueprint-for-closed-source-large-language-models-card.webp","articleSection":["AI and data platforms","Articles","Artificial intelligence","Cross-industry","Generative AI","LLMOps"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint","url":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint","name":"LLMOps platform infrastructure for GenAI solutions","isPartOf":{"@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint#primaryimage"},"image":{"@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint#primaryimage"},"thumbnailUrl":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/05\/llmops-blueprint-for-closed-source-large-language-models-card.webp","datePublished":"2023-10-13T21:12:00+00:00","dateModified":"2025-05-09T10:18:58+00:00","description":"Learn how to migrate GenAI pilot projects to a robust LLMOps platform. Explore technical foundations, guardrails & implementation insights for closed-source LLMs","breadcrumb":{"@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint#primaryimage","url":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/05\/llmops-blueprint-for-closed-source-large-language-models-card.webp","contentUrl":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/05\/llmops-blueprint-for-closed-source-large-language-models-card.webp","width":1496,"height":780,"caption":"LLMOps blueprint for closed-source large language models"},{"@type":"BreadcrumbList","@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/blog\/llmops-platform-blueprint#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/"},{"@type":"ListItem","position":2,"name":"LLMOps blueprint for closed-source large language models"}]},{"@type":"WebSite","@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/#website","url":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/","name":"Grid Dynamics","description":"","publisher":{"@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/#organization","name":"Grid Dynamics","alternateName":"Grid Dynamics Holdings, Inc.","url":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/#\/schema\/logo\/image\/","url":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2023\/10\/share-logo.png","contentUrl":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2023\/10\/share-logo.png","width":720,"height":360,"caption":"Grid Dynamics"},"image":{"@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/#\/schema\/person\/e12e66db4925524135888aea2e3d12eb","name":"Dmitry Mezhensky","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/#\/schema\/person\/image\/","url":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/04\/Dmitry-Mezhensky-150x150.webp","contentUrl":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/04\/Dmitry-Mezhensky-150x150.webp","caption":"Dmitry Mezhensky"},"description":"Dmitry Mezhensky joined Grid Dynamics in 2014 and has worked on various Big Data projects since. One of the major projects, iCrossing, was a huge success as we built a high-performing Big Data platform. Dmitry is currently on-site at a large retailer.","jobTitle":"Director of Big Data and ML Engineering","url":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/author\/dmitry-mezhensky"}]}},"uagb_featured_image_src":{"full":["https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/05\/llmops-blueprint-for-closed-source-large-language-models-card.webp",1496,780,false],"thumbnail":["https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/05\/llmops-blueprint-for-closed-source-large-language-models-card-150x150.webp",150,150,true],"medium":["https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/05\/llmops-blueprint-for-closed-source-large-language-models-card-300x156.webp",300,156,true],"medium_large":["https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/05\/llmops-blueprint-for-closed-source-large-language-models-card-768x400.webp",768,400,true],"large":["https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/05\/llmops-blueprint-for-closed-source-large-language-models-card-1024x534.webp",1024,534,true],"1536x1536":["https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/05\/llmops-blueprint-for-closed-source-large-language-models-card.webp",1496,780,false],"2048x2048":["https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-content\/uploads\/2025\/05\/llmops-blueprint-for-closed-source-large-language-models-card.webp",1496,780,false]},"uagb_author_info":{"display_name":"Dmitry Mezhensky","author_link":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/author\/dmitry-mezhensky"},"uagb_comment_info":0,"uagb_excerpt":"Building solutions using closed-source large language models (LLMs), including models like GPT-4 from OpenAI, or PaLM2 from Google, is a markedly different process to creating private machine learning (ML) models, so traditional MLOps playbooks and best practices might appear irrelevant when applied to LLM-centric projects. And indeed, many companies currently approach LLM projects as greenfield","ai_search_data":{"is_excluded_from_search":false,"custom_thumbnail":"","custom_title":"","custom_label":"","content_type":"Articles"},"_links":{"self":[{"href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-json\/wp\/v2\/posts\/32972","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-json\/wp\/v2\/users\/103"}],"replies":[{"embeddable":true,"href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-json\/wp\/v2\/comments?post=32972"}],"version-history":[{"count":10,"href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-json\/wp\/v2\/posts\/32972\/revisions"}],"predecessor-version":[{"id":63396,"href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-json\/wp\/v2\/posts\/32972\/revisions\/63396"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-json\/wp\/v2\/media\/63225"}],"wp:attachment":[{"href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-json\/wp\/v2\/media?parent=32972"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-json\/wp\/v2\/categories?post=32972"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-json\/wp\/v2\/tags?post=32972"},{"taxonomy":"starter_kit_tag","embeddable":true,"href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/www.griddynamics.com\/wp-json\/wp\/v2\/starter_kit_tag?post=32972"}],"curies":[{"name":"wp","href":"https:\/\/round-lake.dustinice.workers.dev:443\/https\/api.w.org\/{rel}","templated":true}]}}