A safety claim about an AI system is only as credible as the measurement that backs it. Everything else in this book follows from that sentence. If a developer says a model is "safe enough to deploy," a regulator says it is "low-risk," or a buyer says it is "fit for purpose," each of those statements is a load-bearing claim about behavior under conditions that have not yet happened. Without a measurement program - without tests that were specified before the system was built, executed by people who can be honest about the results, and reported in a form that outsiders can interrogate - those claims are aspirations dressed up as findings. This book is about how to tell the difference. AI governance discussions in 2026 are crowded with rules, principles, and voluntary commitments. The substance of those commitments, however, almost always reduces to a claim about what an AI system will or will not do. "We will not deploy a model that meaningfully uplifts the creation of biological weapons." "Our system does not exhibit unacceptable bias in hiring contexts." "The model refuses to generate child sexual abuse material." Each of these claims is a measurement claim. Each requires that someone - the developer, an independent lab, a government body - define what the dangerous behavior looks like, design a probe that can elicit it if it is present, run that probe under conditions representative of real use, and report the result. The NIST AI Risk Management Framework treats this measurement function as one of four continuous functions - Govern, Map, Measure, and Manage (National Institute of Standards and Technology, 2023). The framework does not tell organizations what to measure; it tells them that measurement is a precondition for trustworthy AI rather than a downstream activity. The 2024 Generative AI Profile sharpened this for foundation-model systems, calling out evaluation, documentation, and disclosure controls that have to be in place before generative systems are deployed at scale (National Institute of Standards and Technology, 2024). The framework's posture, read carefully, is that without measurement infrastructure there is no risk management - only assertion. The case for treating measurement as infrastructure, not as ornament, has three parts. First, the systems themselves are now too capable and too widely deployed for narrative assurance to substitute for evidence. External evaluators including METR and its ARC Evals predecessor program have published public materials on frontier capability evaluation and dangerous-capability framing (METR, 2023; METR, 2025); the existence of those public evaluation efforts, whatever one thinks of their methodology, implies that the relevant questions cannot be answered by inspection alone. Second, governments have started to build state evaluation capacity - most visibly the UK AI Security Institute, which has positioned itself as a public-facing evaluation body for advanced AI safety and security (UK AI Security Institute, 2025). Third, the academic and open-source community has produced standing benchmark infrastructure, with Stanford's Holistic Evaluation of Language Models (HELM) project running multi-dimensional evaluations across many models and a broad range of scenarios on a continuing basis (Stanford Center for Research on Foundation Models, 2025). Each of these efforts is partial. Each is contested. Together they constitute the first generation of what it would mean to have actual measurement infrastructure for AI - and they make visible how far that infrastructure still has to go.
Die Inhaltsangabe kann sich auf eine andere Ausgabe dieses Titels beziehen.
Anbieter: Grand Eagle Retail, Bensenville, IL, USA
Paperback. Zustand: new. Paperback. A safety claim about an AI system is only as credible as the measurement that backs it. Everything else in this book follows from that sentence. If a developer says a model is "safe enough to deploy," a regulator says it is "low-risk," or a buyer says it is "fit for purpose," each of those statements is a load-bearing claim about behavior under conditions that have not yet happened. Without a measurement program - without tests that were specified before the system was built, executed by people who can be honest about the results, and reported in a form that outsiders can interrogate - those claims are aspirations dressed up as findings. This book is about how to tell the difference. AI governance discussions in 2026 are crowded with rules, principles, and voluntary commitments. The substance of those commitments, however, almost always reduces to a claim about what an AI system will or will not do. "We will not deploy a model that meaningfully uplifts the creation of biological weapons." "Our system does not exhibit unacceptable bias in hiring contexts." "The model refuses to generate child sexual abuse material." Each of these claims is a measurement claim. Each requires that someone - the developer, an independent lab, a government body - define what the dangerous behavior looks like, design a probe that can elicit it if it is present, run that probe under conditions representative of real use, and report the result. The NIST AI Risk Management Framework treats this measurement function as one of four continuous functions - Govern, Map, Measure, and Manage (National Institute of Standards and Technology, 2023). The framework does not tell organizations what to measure; it tells them that measurement is a precondition for trustworthy AI rather than a downstream activity. The 2024 Generative AI Profile sharpened this for foundation-model systems, calling out evaluation, documentation, and disclosure controls that have to be in place before generative systems are deployed at scale (National Institute of Standards and Technology, 2024). The framework's posture, read carefully, is that without measurement infrastructure there is no risk management - only assertion. The case for treating measurement as infrastructure, not as ornament, has three parts. First, the systems themselves are now too capable and too widely deployed for narrative assurance to substitute for evidence. External evaluators including METR and its ARC Evals predecessor program have published public materials on frontier capability evaluation and dangerous-capability framing (METR, 2023; METR, 2025); the existence of those public evaluation efforts, whatever one thinks of their methodology, implies that the relevant questions cannot be answered by inspection alone. Second, governments have started to build state evaluation capacity - most visibly the UK AI Security Institute, which has positioned itself as a public-facing evaluation body for advanced AI safety and security (UK AI Security Institute, 2025). Third, the academic and open-source community has produced standing benchmark infrastructure, with Stanford's Holistic Evaluation of Language Models (HELM) project running multi-dimensional evaluations across many models and a broad range of scenarios on a continuing basis (Stanford Center for Research on Foundation Models, 2025). Each of these efforts is partial. Each is contested. Together they constitute the first generation of what it would mean to have actual measurement infrastructure for AI - and they make visible how far that infrastructure still has to go. A safety claim about an AI system is only as credible as the measurement that backs it. Everything else in this book follows from that sentence. If a developer says a model is "safe enough to deploy," a regulator says it is "low-risk," or a buyer says it is "fit for purpose," each of those s Shipping may be from multiple locations in the US or from the UK, depending on stock availability. Bestandsnummer des Verkäufers 9798259503489
Anbieter: California Books, Miami, FL, USA
Zustand: New. Bestandsnummer des Verkäufers I-9798259503489
Anzahl: Mehr als 20 verfügbar
Anbieter: PBShop.store US, Wood Dale, IL, USA
PAP. Zustand: New. New Book. Shipped from UK. Established seller since 2000. Bestandsnummer des Verkäufers L2-9798259503489
Anzahl: Mehr als 20 verfügbar
Anbieter: PBShop.store UK, Fairford, GLOS, Vereinigtes Königreich
PAP. Zustand: New. New Book. Shipped from UK. Established seller since 2000. Bestandsnummer des Verkäufers L2-9798259503489
Anzahl: Mehr als 20 verfügbar
Anbieter: AHA-BUCH GmbH, Einbeck, Deutschland
Taschenbuch. Zustand: Neu. nach der Bestellung gedruckt Neuware - Printed after ordering - A safety claim about an AI system is only as credible as the measurement that backs it. Everything else in this book follows from that sentence. If a developer says a model is 'safe enough to deploy,' a regulator says it is 'low-risk,' or a buyer says it is 'fit for purpose,' each of those statements is a load-bearing claim about behavior under . Bestandsnummer des Verkäufers 9798259503489
Anzahl: 2 verfügbar
Anbieter: CitiRetail, Stevenage, Vereinigtes Königreich
Paperback. Zustand: new. Paperback. A safety claim about an AI system is only as credible as the measurement that backs it. Everything else in this book follows from that sentence. If a developer says a model is "safe enough to deploy," a regulator says it is "low-risk," or a buyer says it is "fit for purpose," each of those statements is a load-bearing claim about behavior under conditions that have not yet happened. Without a measurement program - without tests that were specified before the system was built, executed by people who can be honest about the results, and reported in a form that outsiders can interrogate - those claims are aspirations dressed up as findings. This book is about how to tell the difference. AI governance discussions in 2026 are crowded with rules, principles, and voluntary commitments. The substance of those commitments, however, almost always reduces to a claim about what an AI system will or will not do. "We will not deploy a model that meaningfully uplifts the creation of biological weapons." "Our system does not exhibit unacceptable bias in hiring contexts." "The model refuses to generate child sexual abuse material." Each of these claims is a measurement claim. Each requires that someone - the developer, an independent lab, a government body - define what the dangerous behavior looks like, design a probe that can elicit it if it is present, run that probe under conditions representative of real use, and report the result. The NIST AI Risk Management Framework treats this measurement function as one of four continuous functions - Govern, Map, Measure, and Manage (National Institute of Standards and Technology, 2023). The framework does not tell organizations what to measure; it tells them that measurement is a precondition for trustworthy AI rather than a downstream activity. The 2024 Generative AI Profile sharpened this for foundation-model systems, calling out evaluation, documentation, and disclosure controls that have to be in place before generative systems are deployed at scale (National Institute of Standards and Technology, 2024). The framework's posture, read carefully, is that without measurement infrastructure there is no risk management - only assertion. The case for treating measurement as infrastructure, not as ornament, has three parts. First, the systems themselves are now too capable and too widely deployed for narrative assurance to substitute for evidence. External evaluators including METR and its ARC Evals predecessor program have published public materials on frontier capability evaluation and dangerous-capability framing (METR, 2023; METR, 2025); the existence of those public evaluation efforts, whatever one thinks of their methodology, implies that the relevant questions cannot be answered by inspection alone. Second, governments have started to build state evaluation capacity - most visibly the UK AI Security Institute, which has positioned itself as a public-facing evaluation body for advanced AI safety and security (UK AI Security Institute, 2025). Third, the academic and open-source community has produced standing benchmark infrastructure, with Stanford's Holistic Evaluation of Language Models (HELM) project running multi-dimensional evaluations across many models and a broad range of scenarios on a continuing basis (Stanford Center for Research on Foundation Models, 2025). Each of these efforts is partial. Each is contested. Together they constitute the first generation of what it would mean to have actual measurement infrastructure for AI - and they make visible how far that infrastructure still has to go. A safety claim about an AI system is only as credible as the measurement that backs it. Everything else in this book follows from that sentence. If a developer says a model is "safe enough to deploy," a regulator says it is "low-risk," or a buyer says it is "fit for purpose," ea Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability. Bestandsnummer des Verkäufers 9798259503489
Anzahl: 1 verfügbar
Anbieter: preigu, Osnabrück, Deutschland
Taschenbuch. Zustand: Neu. The AI Safety Measurement Problem | What We Can't Test, We Can't Trust | Nimble Books LLC | Taschenbuch | Sprint E | Englisch | 2026 | AI Governance & Strategy | EAN 9798259503489 | Verantwortliche Person für die EU: Libri GmbH, Europaallee 1, 36244 Bad Hersfeld, gpsr[at]libri[dot]de | Anbieter: preigu Print on Demand. Bestandsnummer des Verkäufers 135956825
Anzahl: 5 verfügbar