Isbn: 9781807423896 - practical llm evaluation for production systems: measure, monitor, and improve ai system reliability across training and inference (12 Ergebnisse)

ISBN: 
Mit der Detailsuche verfeinern

Optimieren Sie Ihre Suche

  • Bücher (12)

bis

Benutzerdefinierte Preisspanne (EUR)

bis

  • Sprache: Englisch

    Verlag: Packt Publishing, 2026

    1807423891 / 9781807423896

    • Softcover

    Anbieter: ThriftBooks-Atlanta, AUSTELL, GA, USAThriftBooks-Atlanta

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Gebraucht - Gut

    EUR 38,00

     Versand gratis 
    Versand innerhalb von USA

    Anzahl: 1 verfügbar

    Paperback. Zustand: Very Good. No Jacket. May have limited writing in cover pages. Pages are unmarked. ~ ThriftBooks: Read More, Spend Less.

  • Sprache: Englisch

    Verlag: Packt Publishing 6/30/2026, 2026

    1807423891 / 9781807423896

    • Softcover

    Anbieter: BargainBookStores, Grand Rapids, MI, USABargainBookStores

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 49,51

     Versand gratis 
    Versand innerhalb von USA

    Anzahl: 5 verfügbar

    Paperback or Softback. Zustand: New. Practical LLM Evaluation for Production Systems: Measure, monitor, and improve AI system reliability across training and inference. Book.

  • Sprache: Englisch

    Verlag: Packt Publishing, 2026

    1807423891 / 9781807423896

    • Softcover

    Anbieter: California Books, Miami, FL, USACalifornia Books

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 54,01

     Versand gratis 
    Versand innerhalb von USA

    Anzahl: Mehr als 20 verfügbar

    Zustand: New.

  • Sprache: Englisch

    Verlag: Packt Publishing, 2026

    1807423891 / 9781807423896

    • Softcover

    Anbieter: PBShop.store US, Wood Dale, IL, USAPBShop.store US

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 59,91

     Versand gratis 
    Versand innerhalb von USA

    Anzahl: Mehr als 20 verfügbar

    PAP. Zustand: New. New Book. Shipped from UK. Established seller since 2000.

  • Sprache: Englisch

    Verlag: Packt Publishing, 2026

    1807423891 / 9781807423896

    • Softcover

    Anbieter: PBShop.store UK, Fairford, GLOS, Vereinigtes KönigreichPBShop.store UK

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 53,41

    EUR 6,91 Versand 
    Versand von Vereinigtes Königreich nach USA

    Anzahl: Mehr als 20 verfügbar

    PAP. Zustand: New. New Book. Shipped from UK. Established seller since 2000.

  • Sprache: Englisch

    Verlag: Packt Publishing Limited, GB, 2026

    1807423891 / 9781807423896

    • Softcover

    Anbieter: Rarewaves.com USA, London, LONDO, Vereinigtes KönigreichRarewaves.com USA

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 63,12

     Versand gratis 
    Versand von Vereinigtes Königreich nach USA

    Anzahl: Mehr als 20 verfügbar

    Paperback. Zustand: New. Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applicationsFree with your book: DRM-free PDF version + access to Packt's next-gen Reader*Key FeaturesDesign evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systemsMeasure quality, safety, grounding, robustness, and production readiness with practical metricsApply unified evaluation methods to text, multimodal, and agentic AI systemsBook DescriptionModern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.*Email sign-up and proof of purchase requiredWhat you will learnDesign repeatable evaluation pipelines for LLM systemsAssess inference quality, latency, and operational costEvaluate multimodal, agentic, and reasoning AI systemsBuild regression gates and deployment evaluation workflowsDetect hallucinations and grounding failures in VLMsAssess routing stability in mixture-of-experts modelsEvaluate Text2SQL, OCR, and retrieval-based systemsTranslate evaluation signals into production decisionsWho this book is forML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts.…

  • Sprache: Englisch

    Verlag: Packt Publishing Limited, GB, 2026

    1807423891 / 9781807423896

    • Softcover

    Anbieter: Rarewaves.com UK, London, Vereinigtes KönigreichRarewaves.com UK

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 60,62

    EUR 76,48 Versand 
    Versand von Vereinigtes Königreich nach USA

    Anzahl: Mehr als 20 verfügbar

    Paperback. Zustand: New. Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applicationsFree with your book: DRM-free PDF version + access to Packt's next-gen Reader*Key FeaturesDesign evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systemsMeasure quality, safety, grounding, robustness, and production readiness with practical metricsApply unified evaluation methods to text, multimodal, and agentic AI systemsBook DescriptionModern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.*Email sign-up and proof of purchase requiredWhat you will learnDesign repeatable evaluation pipelines for LLM systemsAssess inference quality, latency, and operational costEvaluate multimodal, agentic, and reasoning AI systemsBuild regression gates and deployment evaluation workflowsDetect hallucinations and grounding failures in VLMsAssess routing stability in mixture-of-experts modelsEvaluate Text2SQL, OCR, and retrieval-based systemsTranslate evaluation signals into production decisionsWho this book is forML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts.…

  • Sprache: Englisch

    Verlag: Packt Publishing Limited, Birmingham, 2026

    1807423891 / 9781807423896

    • Softcover
    • Print-on-Demand

    Anbieter: Grand Eagle Retail, Bensenville, IL, USAGrand Eagle Retail

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 56,33

     Versand gratis 
    Versand innerhalb von USA

    Anzahl: 1 verfügbar

    Paperback. Zustand: new. Paperback. Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applicationsFree with your book: DRM-free PDF version + access to Packt's next-gen Reader*Key FeaturesDesign evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systemsMeasure quality, safety, grounding, robustness, and production readiness with practical metricsApply unified evaluation methods to text, multimodal, and agentic AI systemsBook DescriptionModern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.*Email sign-up and proof of purchase requiredWhat you will learnDesign repeatable evaluation pipelines for LLM systemsAssess inference quality, latency, and operational costEvaluate multimodal, agentic, and reasoning AI systemsBuild regression gates and deployment evaluation workflowsDetect hallucinations and grounding failures in VLMsAssess routing stability in mixture-of-experts modelsEvaluate Text2SQL, OCR, and retrieval-based systemsTranslate evaluation signals into production decisionsWho this book is forML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts. This item is printed on demand. Shipping may be from multiple locations in the US or from the UK, depending on stock availability.…

  • Sprache: Englisch

    Verlag: Packt Publishing Limited, Birmingham, 2026

    1807423891 / 9781807423896

    • Softcover
    • Print-on-Demand

    Anbieter: CitiRetail, Stevenage, Vereinigtes KönigreichCitiRetail

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 58,16

    EUR 43,54 Versand 
    Versand von Vereinigtes Königreich nach USA

    Anzahl: 1 verfügbar

    Paperback. Zustand: new. Paperback. Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applicationsFree with your book: DRM-free PDF version + access to Packt's next-gen Reader*Key FeaturesDesign evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systemsMeasure quality, safety, grounding, robustness, and production readiness with practical metricsApply unified evaluation methods to text, multimodal, and agentic AI systemsBook DescriptionModern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.*Email sign-up and proof of purchase requiredWhat you will learnDesign repeatable evaluation pipelines for LLM systemsAssess inference quality, latency, and operational costEvaluate multimodal, agentic, and reasoning AI systemsBuild regression gates and deployment evaluation workflowsDetect hallucinations and grounding failures in VLMsAssess routing stability in mixture-of-experts modelsEvaluate Text2SQL, OCR, and retrieval-based systemsTranslate evaluation signals into production decisionsWho this book is forML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.…

  • Sprache: Englisch

    Verlag: Packt Publishing, 2026

    1807423891 / 9781807423896

    • Softcover
    • Print-on-Demand

    Anbieter: AHA-BUCH GmbH, Einbeck, DeutschlandAHA-BUCH GmbH

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 71,85

    EUR 35,00 Versand 
    Versand von Deutschland nach USA

    Anzahl: 2 verfügbar

    Taschenbuch. Zustand: Neu. nach der Bestellung gedruckt Neuware - Printed after ordering.

  • Sprache: Englisch

    Verlag: Packt Publishing Limited, Birmingham, 2026

    1807423891 / 9781807423896

    • Softcover
    • Print-on-Demand

    Anbieter: AussieBookSeller, Truganina, VIC, AustralienAussieBookSeller

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 82,45

    EUR 32,88 Versand 
    Versand von Australien nach USA

    Anzahl: 1 verfügbar

    Paperback. Zustand: new. Paperback. Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applicationsFree with your book: DRM-free PDF version + access to Packt's next-gen Reader*Key FeaturesDesign evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systemsMeasure quality, safety, grounding, robustness, and production readiness with practical metricsApply unified evaluation methods to text, multimodal, and agentic AI systemsBook DescriptionModern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.*Email sign-up and proof of purchase requiredWhat you will learnDesign repeatable evaluation pipelines for LLM systemsAssess inference quality, latency, and operational costEvaluate multimodal, agentic, and reasoning AI systemsBuild regression gates and deployment evaluation workflowsDetect hallucinations and grounding failures in VLMsAssess routing stability in mixture-of-experts modelsEvaluate Text2SQL, OCR, and retrieval-based systemsTranslate evaluation signals into production decisionsWho this book is forML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts. This item is printed on demand. Shipping may be from our Sydney, NSW warehouse or from our UK or US warehouse, depending on stock availability.…

  • Sprache: Englisch

    Verlag: Packt Publishing, 2026

    1807423891 / 9781807423896

    • Softcover
    • Print-on-Demand

    Anbieter: preigu, Osnabrück, Deutschlandpreigu

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 65,30

    EUR 70,00 Versand 
    Versand von Deutschland nach USA

    Anzahl: 5 verfügbar

    Taschenbuch. Zustand: Neu. Practical LLM Evaluation for Production Systems | Measure, monitor, and improve AI system reliability across training and inference | Ammar Mohanna (u. a.) | Taschenbuch | Englisch | 2026 | Packt Publishing | EAN 9781807423896 | Verantwortliche Person für die EU: Libri GmbH, Europaallee 1, 36244 Bad Hersfeld, gpsr[at]libri[dot]de | Anbieter: preigu Print on Demand.…