healthcarereimagined

Envisioning healthcare for the 21st century

  • About
  • Economics

AI Report Shows ‘Startlingly Rapid’ Progress—And Ballooning Costs – Scientific American

Posted by timmreardon on 04/26/2024
Posted in: Uncategorized.

A new report finds that AI matches or outperforms people at tasks such as competitive math and reading comprehension

BY NICOLA JONES & NATURE MAGAZINE

Artificial intelligence (AI) systems, such as the chatbot ChatGPT, have become so advanced that they now very nearly match or exceed human performance in tasks including reading comprehension, image classification and competition-level mathematics, according to a new report. Rapid progress in the development of these systems also means that many common benchmarks and tests for assessing them are quickly becoming obsolete.

These are just a few of the top-line findings from the Artificial Intelligence Index Report 2024, which was published on 15 April by the Institute for Human-Centered Artificial Intelligence at Stanford University in California. The report charts the meteoric progress in machine-learning systems over the past decade.

In particular, the report says, new ways of assessing AI — for example, evaluating their performance on complex tasks, such as abstraction and reasoning — are more and more necessary. “A decade ago, benchmarks would serve the community for 5–10 years” whereas now they often become irrelevant in just a few years, says Nestor Maslej, a social scientist at Stanford and editor-in-chief of the AI Index. “The pace of gain has been startlingly rapid.”

Stanford’s annual AI Index, first published in 2017, is compiled by a group of academic and industry specialists to assess the field’s technical capabilities, costs, ethics and more — with an eye towards informing researchers, policymakers and the public. This year’s report, which is more than 400 pages long and was copy-edited and tightened with the aid of AI tools, notes that AI-related regulation in the United States is sharply rising. But the lack of standardized assessments for responsible use of AI makes it difficult to compare systems in terms of the risks that they pose.

The rising use of AI in science is also highlighted in this year’s edition: for the first time, it dedicates an entire chapter to science applications, highlighting projects including Graph Networks for Materials Exploration (GNoME), a project from Google DeepMind that aims to help chemists discover materials, and GraphCast, another DeepMind tool, which does rapid weather forecasting.

GROWING UP

The current AI boom — built on neural networks and machine-learning algorithms — dates back to the early 2010s. The field has since rapidly expanded. For example, the number of AI coding projects on GitHub, a common platform for sharing code, increased from about 800 in 2011 to 1.8 million last year. And journal publications about AI roughly tripled over this period, the report says.

Much of the cutting-edge work on AI is being done in industry: that sector produced 51 notable machine-learning systems last year, whereas academic researchers contributed 15. “Academic work is shifting to analysing the models coming out of companies — doing a deeper dive into their weaknesses,” says Raymond Mooney, director of the AI Lab at the University of Texas at Austin, who wasn’t involved in the report.

That includes developing tougher tests to assess the visual, mathematical and even moral-reasoning capabilities of large language models (LLMs), which power chatbots. One of the latest tests is the Graduate-Level Google-Proof Q&A Benchmark (GPQA), developed last year by a team including machine-learning researcher David Rein at New York University.

The GPQA, consisting of more than 400 multiple-choice questions, is tough: PhD-level scholars could correctly answer questions in their field 65% of the time. The same scholars, when attempting to answer questions outside their field, scored only 34%, despite having access to the Internet during the test (randomly selecting answers would yield a score of 25%). As of last year, AI systems scored about 30–40%. This year, Rein says, Claude 3 — the latest chatbot released by AI company Anthropic, based in San Francisco, California — scored about 60%. “The rate of progress is pretty shocking to a lot of people, me included,” Rein adds. “It’s quite difficult to make a benchmark that survives for more than a few years.”

COST OF BUSINESS

As performance is skyrocketing, so are costs. GPT-4 — the LLM that powers ChatGPT and that was released in March 2023 by San Francisco-based firm OpenAI — reportedly cost US$78 million to train. Google’s chatbot Gemini Ultra, launched in December, cost $191 million. Many people are concerned about the energy use of these systems, as well as the amount of water needed to cool the data centres that help to run them. “These systems are impressive, but they’re also very inefficient,” Maslej says.

Costs and energy use for AI models are high in large part because one of the main ways to make current systems better is to make them bigger. This means training them on ever-larger stocks of text and images. The AI Index notes that some researchers now worry about running out of training data. Last year, according to the report, the non-profit research institute Epoch projected that we might exhaust supplies of high-quality language data as soon as this year. (However, the institute’s most recent analysis suggests that 2028 is a better estimate.)

Ethical concerns about how AI is built and used are also mounting. “People are way more nervous about AI than ever before, both in the United States and across the globe,” says Maslej, who sees signs of a growing international divide. “There are now some countries very excited about AI, and others that are very pessimistic.”

In the United States, the report notes a steep rise in regulatory interest. In 2016, there was just one US regulation that mentioned AI; last year, there were 25. “After 2022, there’s a massive spike in the number of AI-related bills that have been proposed” by policymakers, Maslej says.

Regulatory action is increasingly focused on promoting responsible AI use. Although benchmarks are emerging that can score metrics such as an AI tool’s truthfulness, bias and even likability, not everyone is using the same models, Maslej says, which makes cross-comparisons hard. “This is a really important topic,” he says. “We need to bring the community together on this.”

This article is reproduced with permission and was first published on April 15, 2024.

Article link: https://www.scientificamerican.com/article/stanford-ai-index-rapid-progress/?

Advanced SAM.gov | Understanding Notice Types

Posted by timmreardon on 04/23/2024
Posted in: Uncategorized.

The main tool federal agencies use to communicate with government contractors is SAM.gov. Become an advance user of the buyers’ communication tool to win more.

This article summarizes my slides from a training I did today (4/22/2024) about the 9 Notice Types in SAM.

Understand SAM’s Fit in Federal Buyer Toolbox

Federal agency buyers communicate with government contractors (industry) using multiple tools. Here are a few of the most important ones:

  • System for Award Management (SAM)
  • Federal Procurement Data System (FPDS)
  • USASpending | Primarily a nice dashboard of FPDS data
  • Contractor Performance Assessment Reporting System (CPARS) | This is where federal buyers communicate to each other and to incumbents about contract performance.
  • Bid Boards (e.g., DIBBS, ARC, etc.) | DIBBS is a DLA based tool primarily for product purchases. ARC is the tool used by the Intelligence Community primarily since there are classification concerns.
  • Contract Vehicle Portals | For example, Seaport NxG is the Navy’s primary contract vehicle. Notices appearing here, do not appear in SAM.

The 9 Notice Types in SAM

  1. Special Notice | Heads up on Industry Days, APBI, or LRAFs
  2. Sources Sought | Seeking possible vendors
  3. Presolicitation | Makes known a solicitation may follow soon
  4. Consolidate / (substantially) Bundle | Intent to Bundle Requirements
  5. Solicitation | RFQ, RFP, etc.
  6. Combined Synopsis / Solicitation | Solicitations with specifications
  7. Award Notice | Vendor who received an award and the amount agreed
  8. Justification | Needed when a solicitation isn’t posted and one vendor used
  9. Sale of Surplus Property | Usually for real estate no longer needed

Importance of Each Notice Type

I racked and stacked the notice types based on their value to your company. The rationale I used is ‘shift left’ – any notice that helps you learn about and engage on an opportunity sooner is better.

Low Value

  • Award Notice
  • Consolidate / (substantially) Bundle
  • Justification
  • Sale of Surplus Property

Medium Value

  • Combined Synopsis / Solicitation
  • Solicitation

High Value

  • Sources Sought
  • Special Notices
  • Presolicitation

How to Rapidly Process Many SAM.Gov Notices

With hundreds or even thousands of opportunities being processed in SAM, you need a way to process them fast to find those best for you.

Before you follow my three-step approach below, make sure you document your ‘Slam Dunk’ opportunity criteria. Watch my other training to understand how to define your ‘slam dunk’ criteria in 30 minutes or less.

Three Step Approach

  1. Triage By Title | Seconds Per Opportunity | Use Tabs
  2. Triage By Description | SAM Field Only
  3. Triage By Files | Before Adding to Pipeline

Video Replay of Notice Types in SAM.Gov

Here’s a replay from the training I did here on LinkedIn about the 9 types of notice within the System for Awards Management (SAM).

View media

Article link: https://www.linkedin.com/pulse/advanced-samgov-understanding-notice-types-neil-mcdonnell-gnsee

The who, what, and where of AI adoption in America – MIT Sloan

Posted by timmreardon on 04/20/2024
Posted in: Uncategorized.

by Brian Eastwood

Feb 7, 2024

Why It Matters

A new study finds that artificial intelligence is being adopted unevenly in the U.S., with use clustered in large companies and industries such as manufacturing and health care.

It’s not hard to find headlines that suggest artificial intelligence is taking over the business world, from content creation to decision support and process automation.

But reality looks different. A new working paper from the National Bureau of Economic Research about early adoption of AI in the U.S. provides a more nuanced look at which companies are adopting AI, where they are located, and what technologies they are using.

The research shows variation in AI adoption, according to Kristina McElheran, a visiting scholar with the MIT Initiative on the Digital Economy and the paper’s lead author. Just 6% of U.S. companies used AI in 2017, the researchers found, and AI use was concentrated in larger companies and in industries such as manufacturing and information technology. Adoption was also clustered in some “superstar” cities, such as San Francisco, San Antonio, and Nashville.

“The narrative is that AI is everywhere all at once, but the data shows it’s harder to do than people seem interested in discussing,” said McElheran, an assistant professor at the University of Toronto.

“The digital age has arrived, but it has arrived unevenly,” she said.  

AI use in America: large companies, certain sectors 

Research about AI adoption tends to focus on indirect measures of economic activity that refer to AI use — patents, academic publications, or job descriptions that mention AI, McElheran said.

For a more direct measurement, the researchers joined forces with the U.S. Census Bureau and the National Center for Science and Engineering Statistics to conduct the newly developed Annual Business Survey beginning in 2018. The survey asked firms to describe their use of digital information, cloud computing, types of AI, and other advanced technologies in the prior year. The researchers took data from 447,000 responses from the 2018 survey, linked it to 2017 data in the Census Bureau’s Longitudinal Business Database, and weighted it to represent more than 4 million firms nationwide.

The researchers defined AI adoption as using AI for production — “not in invention, not in aspiration, and not even in commercialization from firms that are selling things that rely on AI,” McElheran said.

The finding that just 6% of companies reported using AI in 2017 is still relevant today, McElheran said,  pointing to a November 2023 Census Bureau survey that showed that fewer than 4% of companies use AI to produce goods and services.

The initial, in-depth survey showed other early trends:

  • AI use was highest among large companies. More than 50% of companies with more than 5,000 employees were using AI, as were more than 60% of companies with more than 10,000 employees.
  • Use varied among sectors. About 12% of firms in manufacturing, information services, and health care were using AI, compared with 4% in construction and retail.
  • AI adoption is happening in some superstar cities, but it has also clustered in some unlikely places. These include manufacturing hubs in the Midwest as well as Southern cities with fewer companies overall than tech hubs in Silicon Valley, the Boston area, or New York City. “Use of AI in production is happening in different places than just the areas that are inventing and commercializing AI-based technologies,” McElheran said.

Startups that embrace AI have younger leaders 

To help determine the characteristics of companies that are more likely to use AI, the researchers identified 75,000 startups that participated in the 2018 Annual Business Survey and weighted their responses to represent 740,000 firms.

The researchers found that startups using AI were more likely to have younger, more highly educated, and more highly experienced leaders than startups that were not using AI. Venture capital backing and a focus on process innovation were also associated with AI adoption.  

“The firms that have other things going for them tend to be the ones that can leverage bleeding-edge technology like AI,” McElheran said. “The ability to reconfigure how work gets done and how things get made is an important predictor of whether AI is used in production.”

This matters when comparing AI to other types of general-purpose technologies. Innovations such as enterprise software are complex implementations that depend on a completely different set of workflows.

But AI is more similar to a point solution, McElheran said. “At an incremental level, you can transform a given task, or replicate an individual human task,” she said. “It’s not suddenly everywhere all at once.”

This blessing can quickly become a curse, though. Innovate one part of a system, McElheran noted, and the rest of the system needs to innovate at the same pace. Otherwise, “things start to come unglued.” That’s why firms focusing on process innovation — and benefiting from the resources necessary to move process innovation along — are more likely than others to be using AI.

Some of those AI users are in sectors not typically associated with cutting-edge technology, such as manufacturing and health care. The former is closely linked to manufacturing’s use of robotics. The latter stems from a range of use cases, from optimizing operating room schedules to automating back-office coding and billing processes.

AI adoption requires overcoming inertia and adjustment costs

Ultimately, the biggest barriers to AI adoption may be inertia and adjustment costs. This was true with the internet, word processors, and even double-entry bookkeeping. Both factors exist for good reasons, McElheran said, and shouldn’t be discounted.

Routine is embedded in work practices at many companies. “What do you do at the office every Monday morning?” McElheran asked. “Very few people start with a blank slate to redesign the activities that occupy their time and attention. For reasons we’ve known since the steam engine, it takes a while for firms and for people to adjust.” While helpful for day-to-day operations, routines tend to work against change.

Adopting new technology also typical entails costs somewhere. Firms are primed to use it, and consumers are primed to benefit from it, but these gains don’t come for free. Competition can lead to job losses and other economic adjustments. Firms prioritize workers who already possess the skills to use new tech. As noted in another paper co-authored by McElheran, this means workers over age 50 often miss out on the same salary increases their younger colleagues enjoy from digital transformation.

“When we see trends with upside potential, we can’t ignore the dark side that can overturn the aspirations that people have for their jobs and their children” McElheran said. “We need an approach to AI that is realistic and evidence-based about both the benefits and costs for different pockets of the economy and society.”

The paper is authored by McElheran; University of British Columbia professor J. Frank Li; Stanford University professor Erik Brynjolfsson, PhD ’91; and U.S. Census Bureau economists Zachary Kroff, Emin Dinlersoz, Lucia S. Foster, and Nikolas Zolas.

Article link: https://mitsloan.mit.edu/ideas-made-to-matter/who-what-and-where-ai-adoption-america?

TSMC’s stalled Arizona chip factory is ‘well on track’ to start production next year — and it’ll be charging more for US-made chips – Business Insider

Posted by timmreardon on 04/19/2024
Posted in: Uncategorized.

Jacob Zinkula 

Apr 19, 2024, 6:03 AM EDT

  • TSMC’s Arizona chip factories have faced construction delays. 
  • But the company said it’s “well on track” to start producing chips at its first factory in 2025.
  • TSMC plans to charge more for chips made outside Taiwan to combat higher manufacturing costs. 

Things may be starting to look up for the world’s leading chipmaker.

Last year, Taiwan Semiconductor Manufacturing Company reported its first profit decline in four years. But on April 18, the company reported its strongest sales growth since 2022, and rising quarterly profits that beat expectations. The Taiwan-based TSMC also forecast that second-quarter sales could rise as much as 30% on the backs of “insatiable” demand for chips used to power AI technologies like ChatGPT.

But for the US, in particular, the most important detail from the call may have been the update on the construction timeline of TSMC’s Arizona chips factories. TSMC said it had made “significant progress” on the construction of its first Arizona factory — located in the Phoenix area — and that it was “well on track” to begin producing chips in the first half of 2025. The company said engineering wafer production began at the factory in April, an important step toward the eventual chip production.

The chipmaker’s commitment to building three factories on its Phoenix campus is a key pillar of the Biden administration’s efforts to boost the US’s manufacturing of chips that power everything from cars to iPhones. Bolstering domestic manufacturing could also make the US less reliant on Taiwan — which faces the potential risk of a Chinese invasion.

TSMC’s progress is also important for President Joe Biden because Arizona is a key swing state in the upcoming presidential election. The company’s investment is expected to create roughly 6,000 “high wage” jobs across the factories, in addition to over 20,000 construction jobs, and tens of thousands of indirect supplier jobs.

However, construction has faced a series of challenges. Last July, TSMC announced that chip production for the first factory would be postponed from 2024 to 2025. A lack of skilled construction workers in the US was cited as a reason for the first factory’s delay. Additionally, in January, the opening of its second factory was delayed from 2026 to 2027 or 2028.

Barring further setbacks, TSMC’s update could mean the first factory will begin production of chips in 2025. In recent weeks, however, a report from the Chinese news outlet money.udn has fed speculation among some experts that production could begin by the end of 2024 — TSMC has stuck to the 2025 timeline in public comments.

The sooner chip production begins, the sooner Americans will have access to the “long term,” non-construction jobs TSMC has promised, Dylan Patel, a chief analyst at the semiconductor research and consulting firm SemiAnalysis, told Business Insider.

During the earnings call, TSMC said 2028 was the scheduled opening of the second factory. The third factory is expected to begin production by 2030.

TSMC is planning to charge more for chips made outside Taiwan

Earlier this month, TSMC got more good news: The Biden administration announced it was providing the company with up to $6.6 billion in direct funding and an additional $5 billion in proposed loans to support its investment in Arizona.

Chipmakers have been vying for funding from the CHIPS and Science Act, legislation passed in 2022 that’s expected to fund over $200 billion in US chip production.

This funding could be particularly important for TSMC, given the cost of factory construction and chip manufacturing can differ betweenthe US and Taiwan.

In 2022, TSMC’s founder Morris Chang said that US efforts to boost chip production would be “a wasteful, expensive exercise in futility,” adding that “manufacturing chips in the US is 50% more expensive than in Taiwan.”

In its first-quarter earnings call, TSMC said that cost pressures would cause it to charge more for chips made outside Taiwan, the Financial Times reported. The company also has plans to buildtwo factories in Japan and one in Germany.

“If a customer requests to be in a certain geographical area, the customer needs to share the incremental cost,” TSMC CEO C.C. Wei said during the earnings call.

While boosting the US manufacturing of chips and other products could create jobs and help secure supply chains, it could also lead to higher prices for American consumers.

If Apple, for instance, follows through on its commitment to source chips from TSMC’s Arizona factories, it could make the latest iPhone more expensive.

Article link: https://www.businessinsider.com/tsmc-arizona-semiconductor-chip-fab-taiwan-china-president-joe-biden-2024-4

The dust has settled from the AI executive order – Here’s what agencies should tackle next – Federal News Network

Posted by timmreardon on 04/18/2024
Posted in: Uncategorized.

While it’s clear the government has made progress since the initial guidance was issued, there’s still much to be done to support overall safe federal AI.

Gaurav Pal

April 3, 2024 10:31 am

fter the dust has settled around the much anticipated AI executive order, the White House recently released a fact sheetannouncing key actions as a follow-up three months later. The document summarizes actions that agencies have taken since the EO was issued, including highlights to managing risks and safety measures and investments into innovation.

While it’s clear the government has been making progress since the initial guidance was issued, there’s still much to be done to support overall safe federal AI adoption, including prioritizing security and standardizing guidance. To accomplish this undertaking, federal agencies can look to existing frameworks and resources and apply them to artificial intelligence to accelerate safe AI adoption.

It’s no longer a question of if AI is going to be implemented across the federal government – it’s a question of how, and how fast can it be implemented in a secure manner?

Progress made since the AI EO release

Implementing AI across the federal government has been a massive undertaking, with many agencies starting at ground zero at the start of last year. Since then, the White House has made it clear that implementing AI in a safe and ethical manner is a key priority for the administration, issuing major guidance and directives over the past several months.

According to the AI EO follow-up fact sheet, key targets have been hit in several areas including:

  • Managing risks to safety and security: Completed risk assessments covering AI’s use in every critical infrastructure sector are the most crucial area.
  • Innovating AI for good: Included launches of several AI pilots, research and funding initiatives across key focus areas including HHS and K-12 education.

What should agencies tackle next?

Agencies should further lean into safety and security considerations to ensure AI is being used responsibly and in a manner that protects agencies’ critical data and resources. In January, the National Institute of Standards and Technology released a publication warning regarding privacy and security challenges arising from rapid AI deployment. The publication urges that security needs to be of the utmost importance for any public sector agency interested in implementing AI, which should be the next priority agencies tackle along their AI journeys.

Looking back on similar major technology transformations over the past couple years, such as cloud migration, we can begin to understand what the current problems are. It took the federal government over a decade to really nail down the details of ensuring cloud technology was secure — as a result of the federal government’s migration to the cloud, the government released the Federal Risk and Authorization Management Program (FedRAMP) as a form of guidance.

The good news is, we can learn from the lessons of the last ten years of cloud migration to accelerate AI and deliver it faster to the federal government and the American people by extending and leveraging existing governance models including the Federal Information and Security Management Act and FedRAMP Authority to Operate (ATO) by creating overlays for AI-specific safety, bias and explainability risks. ATO is a concept first developed by NIST to create strong governance for IT systems. This concept, along with others, can be applied to AI systems so agencies don’t need to reinvent the wheel when it comes to securing AI and deploying safe systems into production.

Where to get help?

There’s an abundance of trustworthy resources federal leaders can look to for additional guidance. One new initiative to keep an eye on is from NIST’s recently created AI Safety Institute Consortium (AISIC).

AISIC brings together more than 200 leading stakeholders, including AI creators and users, academics, government and industry researchers, and civil society organizations. AISIC’s mission is to develop guidelines and standards for AI measurement and policy, to help our country be prepared for AI adoption with the appropriate risk management strategies needed.

Additionally, agency leaders can look to industry partners with established centers of excellence or advisory committees with cross-sector expertise and third-party validation. Seek out counsel from industry partners that have experience working with or alongside the federal government, that truly understand the challenges that the government faces. The federal government shouldn’t have to go on this journey alone. There are several established working groups and trusted industry partners eager to share their knowledge.

Agencies across a wide range of sectors are continuing to make progress in their AI journeys, and the federal government continues to prioritize implementation guidance. It can be overwhelming to cut through the noise when it comes to what’s truly necessary to consider or to decide what factors to prioritize the most.

Leaders across the federal government must continue to prioritize security, and the best way to do this is by leaning into already published guidelines and seeking the best external resources available. While the federal government works on standardizing guidelines for AI, agencies can have peace of mind by following the roadmaps that they are most familiar with when it comes to best security practices and apply these to artificial intelligence adoption.

Gaurav “GP” Pal is found and CEO of stackArmor.

Article link: https://federalnewsnetwork.com/commentary/2024/04/the-dust-has-settled-from-the-ai-executive-order-heres-what-agencies-should-tackle-next/

The global chip industry’s complicated contours are decades in the making – WEF

Posted by timmreardon on 04/18/2024
Posted in: Uncategorized.

Sep 6, 2023

John Letzing

Digital Editor, World Economic Forum

  • Semiconductors are the lifeblood of economic growth and innovation in fields like artificial intelligence.
  • But the global industry has been shaped in ways that expose it to geopolitical risk.
  • Chris Miller, the author of ‘Chip War,’ spoke with the World Economic Forum’s Radio Davos podcast about the industry’s past and possible future.
  • Subscribe to Radio Davos on any podcast app: https://pod.link/1504682164; or visit wef.ch/podcasts.

The Netherlands has been an innovation engine for centuries, giving us the world’s first multinational corporation, telescope, and cassette tape. Now, it’s an essential link in a backbone of innovative silicon keeping the global economy upright.

The mixture of happenstance and geostrategy that helped make this small European country key to the global semiconductor market is depicted in delicious detail in Chris Miller’s “Chip War.” The book, published last fall, couldn’t have been better timed.

Miller, an associate professor of international history at the Fletcher School, uses a colorful cast of characters to tell the story of a truly pivotal industry’s formation, and explain why altering it in a meaningful way seems unlikely any time soon – regardless of mountinggeopolitical pressure. 

Chips are coveted not least for the role they play in artificial intelligence tools seemingly poised to shake things up for just about everyone. The more we want them, though, the more expensive and difficult they are to make. It’s all gotten very complicated. 

Take the Dutch niche in the supply chain, for example – it’s based on one company’s machine “that took tens of billions of dollars and several decades to develop,” Miller writes, and uses light to print patterns on silicon by deploying lasers that can hit 50 million tin drops per second.

It’s an industry full of such mind-bending extremes. 

In an interview with the Forum’s Radio Davos podcast, Miller marvelled at having recently visited a facility in the US being built with “seventh-biggest crane that exists in the world,” which will eventually assemble chips mounted with transistors “roughly the size of a coronavirus.”

Nvidia, the company now most closely identified with chips powering artificial intelligence, features prominently in Miller’s book. The company traces its roots to a meeting at a 24-hour diner on the fringes of Silicon Valley, he writes. At a certain point it realized that its semiconductors used for video-game graphics could do a good job of training AI systems. Earlier this year, its market value increased by $184 billion in a single day. 

Nvidia’s chips aren’t made anywhere near Silicon Valley, though. Like most advanced semiconductors they’re producedby another company, TSMC, at a facility in Taiwan, China that Miller describes as “most expensive factory in the world.” 

In fact, US chip production in general has declined sharply in recent decades. 

Instead, the country has focused on research and design, while relying on links with East Asia and the Netherlands for other elements. But those links risk becoming “choke points,” as Miller describes them, if they’re disrupted by conflict or a natural disaster (it’s not just the plot of a 1980s James Bond film; Miller noted in his Radio Davos interview that an unsettling amount of the industry is located in places relatively prone to earthquakes). 

These hazards, and global competition that’s formed harder edges of late, have fueled efforts to build chip resilience through greater independence. 

The ongoing race to gain an edge in chips 

That massive crane Miller mentioned is being put to work in the state of Arizona, which may be a key part of a current US government effort to “win the race for the 21st century” through semiconductor manufacturing. 

The EU has its own initiativedesigned to strengthen chip competitiveness and resilience.

And a proposed, $20 billion effort to build India’s first semiconductor factory (or “fab,” in industry lingo) recently fell through when a key partner backed out. 

In his Radio Davos interview, Miller said the daunting size of the previously planned investment in India is about standard for any new, fully-fledged manufacturing facility. Critics of what the US spends on its military might like to know that “making semiconductors is so expensive that even the Pentagon can’t afford to do it in-house,” he writes. 

Sharing the considerable financial burden of making chips was long ago deemed necessary. Research in one country, building elaborate lithography tools in another, manufacturing in another, and finally assembling in yet another. The system works in good times; in less-good times it seems problematic.

The shape of the industry was no accident, though.

“Microelectronics is a mechanical brain,” Soviet leader Nikita Khrushchev pronounced in the depths of the Cold War, according to Miller’s book, “It is our future.” Khrushchev was right, but maybe not exactly in the way he would have liked. 

At that time, the US was only about four years ahead of the Soviets in chip technology, as the industry’s earliest companies like Fairchild Semiconductor and Texas Instruments focused on space exploration and nuclear weapons. 

Once those firms tapped into the vast American consumer market via electronics, the rest was history, Miller writes. An arms race with nuclear warheads was one thing, a race to cram millions of transistors onto a single chip was another. The Soviets fell behind, and Asia came to the fore. 

Fairchild began sending its chips to Hong Kong SAR for assembly in the early 1960s. A couple of decades after that, a one-time English literature student named Morris Chang founded TSMC in Taiwan, China. The company now churns out roughly 90% of the world’s advanced chips, and has recently been filing a sizeable portion of global semiconductor patent applications. 

Having the right chips or not can make a big difference in a technology market, or on a battlefield. 

But, as Miller notes, going it alone in such an expensive and complex industry has never worked. It’s unclear whether forming distinct, competing supply chains would be much better.

One of the most compelling points Miller makes is that among the many things about chips we take for granted, the biggest might be the mind-blowing increases in computing power they give us year after year. 

But there’s no guarantee that will continue. Moore’s Law, which long ago posited that the power crammed onto a single chip would double about every two years, has so far proven resilient. But it isn’t really a law – it’s just an educated guess.

More reading on chips and global competition 

For more context, here are links to further reading from the World Economic Forum’s Strategic Intelligence platform:

  • Making chips is “an almost incomprehensibly precise, difficult and expensive business,” according to this piece. That means greater collaboration will be essential. (Scientific American)
  • “There were challenging gaps we were not able to smoothly overcome.” This piece reports on what would’ve been India’s first big-ticket semiconductor fab. (The Diplomat)
  • The fact that governments are spending bigger amounts to subsidize domestic chip industries promises unpredictable global consequences, according to this analysis. (Lowy Institute)
  • Morris Chang and the “silicon shield” – this piece digs into the geopolitical fault line running through the “most indispensable” economy in the world (naturally, it draws on Chris Miller’s book). (The Conversation)
  • “The US has pursued its semiconductor strategy without leveraging its greatest strength: its allies.” This piece argues that the country should appoint a special envoy for chips. (The Diplomat)
  • An alternative to silicon for powering the 7G networks of the future? Transistors eventually won’t be able to get any smaller, so this research delves into ways of using new types of materials to make them smarter instead. (Science Daily)
  • Semiconductor export controls may be a precursor to what’s yet to come with quantum computing, which according to this piece is the next emerging technology stirring fears of weaponization. (Harvard Kennedy School)

On the Strategic Intelligenceplatform, you can find feeds of expert analysis related to the Future of Computing, Trade, Geopolitics and hundreds of additional topics. You’ll need to register to view.

Article link: https://www.weforum.org/agenda/2023/09/the-global-chip-industrys-complicated-contours-were-decades-in-the-making/?

The law aims to ensure large AI models don’t pose risks to democracy – WEF

Posted by timmreardon on 04/17/2024
Posted in: Uncategorized.

Learn more from the Forum’s briefing papers on the responsible development of artificial intelligence: https://ow.ly/uyky50RarsP

European Commission Eva Maydell (Paunova)

Reports

Published: 18 January 2024

AI Governance Alliance: Briefing Paper Series

Download PDF

In an era marked by rapid technological transformation, this briefing paper series stands as a pivotal point of reference, guiding responsible transformation with artificial intelligence (AI).

In an era marked by rapid technological transformation, this briefing paper series stands as a pivotal point of reference, guiding responsible transformation with artificial intelligence (AI). 

This collaborative effort brings together over 250 members from more than 200 organizations. Structured around three core working groups (Safe Systems and Technologies, Responsible Applications and Transformation, and Resilient Governance and Regulation), the AI Governance Alliance addresses AI’s multifaceted challenges and opportunities.

This briefing paper series, representing collective insights, establishes foundational focus areas for steering AI’s development, adoption and governance. The alliance serves as a beacon of multistakeholder collaboration, guiding decision-makers towards an AI future that upholds human values and enhances societal progress.

Paper 1 – Presidio AI Framework: Towards Safe Generative AI Models

This paper navigates the complex rise of generative AI, emphasizing the balance between innovation, safety and ethics. It introduces a comprehensive framework centred on an expanded AI life cycle, robust risk guardrails and a shift-left methodology for early safety integration. Advocating for multistakeholder collaboration, the framework promotes shared responsibility and proactive risk management. 

This foundational paper by the AI Governance Alliance sets the stage for ongoing efforts to ensure ethical and responsible AI development, advocating for a future where innovation is coupled with stringent safety measures.

Read the full report here.

Paper 2 – Unlocking Value from Generative AI: Guidance for Responsible Transformation

This paper examines the disruptive potential of generative AI and the imperative for leaders to adopt a use-case-based approach for its deployment. It guides organizations to assess use cases for business impact, operational readiness and investment strategy, and to balance benefits against potential workforce impact and downstream implications. 

Emphasizing a multistakeholder approach, the paper advocates for responsible scaling strategies like transparent governance and value-based change management. This paper equips leaders with insights to responsibly harness generative AI’s benefits while preparing for its evolving future.

Read the full report here.

Paper 3 – Generative AI Governance: Shaping a Collective Global Future

This paper navigates the complexities of AI governance amidst rapid technological and societal changes. It compares national responses, focusing on governance approaches and regulatory instruments. The paper highlights key debates in generative AI, including risk prioritization and access spectrum, and advocates for international cooperation to prevent governance fragmentation. It emphasizes the need for equitable access and inclusion, especially for the Global South. 

This briefing paper informs stakeholders in AI governance and regulation and lays the groundwork for the World Economic Forum’s AI Governance Alliance’s future initiatives on resilient and inclusive governance.

Read the full report here.

Download PDF

Article link: https://www.linkedin.com/posts/world-economic-forum_the-law-aims-to-ensure-large-ai-models-dont-activity-7184070214290952192-hxGe?

Emerging Technology and Risk Analysis – RAND

Posted by timmreardon on 04/17/2024
Posted in: Uncategorized.

Artificial Intelligence and Critical Infrastructure

Published Apr 2, 2024

by Daniel M. Gerstein, Erin N. Leidy

  • Related Topics: 
  • Artificial Intelligence, 
  • Cybersecurity, 
  • Emerging Technologies, 
  • Homeland Security and Public Safety
  • Citation
  • Synopsis(print-friendly)
DOWNLOAD FREE ELECTRONIC DOCUMENT

PDF file 0.3 MB

Research Questions

  1. What is the technology availability for AI applications in critical infrastructure in the next ten years?
  2. How will science and technology maturity; use case, demand, and market forces; resources; policy, legal, ethical, and regulatory impediments; and technology accessibility of critical infrastructure applications change during this ten-year period?
  3. What risks and scenarios (consisting of threats, vulnerabilities, and consequences) is AI likely to present for critical infrastructure applications in the next ten years?

This report is one in a series of analyses on the effects of emerging technologies on U.S. Department of Homeland Security (DHS) missions and capabilities. As part of this research, the authors were charged with developing a technology and risk assessment methodology for evaluating emerging technologies and understanding their implications within a homeland security context. The methodology and analyses provide a basis for DHS to better understand the emerging technologies and the risks they present.

This report focuses on artificial intelligence (AI), especially as it relates to critical infrastructure. The authors draw on the literature about smart cities and consider four attributes in assessing the technology: technology availability and risks and scenarios (which the authors divided into threat, vulnerability, and consequence). The risks and scenarios considered in this analysis pertain to AI use affecting critical infrastructure. The use cases could be either for monitoring and controlling critical infrastructure or for adversaries employing AI for use in illicit activities and nefarious acts directed at critical infrastructure. The risks and scenarios were provided by the DHS Science and Technology Directorate and the DHS Office of Policy. The authors compared these four attributes across three periods: short term (up to three years), medium term (three to five years), and long term (five to ten years) to assess the availability of and risks associated with AI-enabled critical infrastructure.

Key Findings

  • AI is transformative technology and will likely be incorporated broadly across society—including in critical infrastructure.
  • AI will likely be affected by many of the same factors as other information age technologies, such as cybersecurity, protecting intellectual property, ensuring key data protections, and protecting proprietary methods and processes.
  • The AI field contains numerous technologies that will be incorporated into AI systems as they become available. As a result, AI science and technology maturity will be based on key dependencies in several essential technology areas, including high-performance computing, advanced semiconductor development and manufacturing, robotics, machine learning, natural language processing, and the ability to accumulate and protect key data.
  • To place AI in its current state of maturity, it is useful to delineate three AI categories: artificial narrow intelligence (ANI), artificial general intelligence, and artificial super intelligence. By the end of the ten-year period of this analysis, the technology will very likely still only have achieved ANI.
  • AI will present both opportunities and challenges for critical infrastructure and the eventual development of purpose-built smart cities.
  • The ChatGPT-4 rollout in March 2023 provides an interesting case study for how these AI technologies—in this case, large-language models—are likely to mature and be integrated into society. The initial rollout illustrated a cycle—development, deployment, identification of shortcomings and other areas of potential use, and rapid updating of AI systems—that will likely be a feature of AI.

Article link: https://www.rand.org/pubs/research_reports/RRA2873-1.html?

AI hype is built on high test scores. Those tests are flawed -MIT Technology Review

Posted by timmreardon on 04/11/2024
Posted in: Uncategorized.


With hopes and fears about the technology running wild, it’s time to agree on what it can and can’t do.

By Will Douglas Heaven

August 30, 2023

When Taylor Webb played around with GPT-3 in early 2022, he was blown away by what OpenAI’s large language model appeared to be able to do. Here was a neural network trained only to predict the next word in a block of text—a jumped-up autocomplete. And yet it gave correct answers to many of the abstract problems that Webb set for it—the kind of thing you’d find in an IQ test. “I was really shocked by its ability to solve these problems,” he says. “It completely upended everything I would have predicted.”

Webb is a psychologist at the University of California, Los Angeles, who studies the different ways people and computers solve abstract problems. He was used to building neural networks that had specific reasoning capabilities bolted on. But GPT-3 seemed to have learned them for free.

Last month Webb and his colleagues published an article in Nature, in which they describe GPT-3’s ability to pass a variety of tests devised to assess the use of analogy to solve problems (known as analogical reasoning). On some of those tests GPT-3 scored better than a group of undergrads. “Analogy is central to human reasoning,” says Webb. “We think of it as being one of the major things that any kind of machine intelligence would need to demonstrate.”

What Webb’s research highlights is only the latest in a long string of remarkable tricks pulled off by large language models. For example, when OpenAI unveiled GPT-3’s successor, GPT-4, in March, the company published an eye-popping list of professional and academic assessments that it claimed its new large language model had aced, including a couple of dozen high school tests and the bar exam. OpenAI later worked with Microsoft to show that GPT-4 could pass parts of the United States Medical Licensing Examination.

And multiple researchers claim to have shown that large language models can pass tests designed to identify certain cognitive abilities in humans, from chain-of-thought reasoning (working through a problem step by step) to theory of mind (guessing what other people are thinking). 

Such results are feeding a hype machine that predicts computers will soon come for white-collar jobs, replacing teachers, journalists, lawyers and more. Geoffrey Hinton has called out GPT-4’s apparent ability to string together thoughts as one reason he is now scared of the technology he helped create. 

But there’s a problem: there is little agreement on what those results actually mean. Some people are dazzled by what they see as glimmers of human-like intelligence. Others aren’t convinced one bit.

“There are several critical issues with current evaluation techniques for large language models,” says Natalie Shapira, a computer scientist at Bar-Ilan University in Ramat Gan, Israel. “It creates the illusion that they have greater capabilities than what truly exists.”

That’s why a growing number of researchers—computer scientists, cognitive scientists, neuroscientists, linguists—want to overhaul the way large language models are assessed, calling for more rigorous and exhaustive evaluation. Some think that the practice of scoring machines on human tests is wrongheaded, period, and should be ditched.

“People have been giving human intelligence tests—IQ tests and so on—to machines since the very beginning of AI,” says Melanie Mitchell, an artificial-intelligence researcher at the Santa Fe Institute in New Mexico. “The issue throughout has been what it means when you test a machine like this. It doesn’t mean the same thing that it means for a human.”

“There’s a lot of anthropomorphizing going on,” she says. “And that’s kind of coloring the way that we think about these systems and how we test them.”

With hopes and fears for this technology at an all-time high, it is crucial that we get a solid grip on what large language models can and cannot do. 

Open to interpretation

Most of the problems with testing large language models boil down to the question of how to interpret the results. 

Assessments designed for humans, like high school exams and IQ tests, take a lot for granted. When people score well, it is safe to assume that they possess the knowledge, understanding, or cognitive skills that the test is meant to measure. (In practice, that assumption only goes so far. Academic exams do not always reflect students’ true abilities. IQ tests measure a specific set of skills, not overall intelligence. Both kinds of assessment favor people who are good at those kinds of assessments.) 

But when a large language model scores well on such tests, it is not clear at all what has been measured. Is it evidence of actual understanding? A mindless statistical trick? Rote repetition?

“There is a long history of developing methods to test the human mind,” says Laura Weidinger, a senior research scientist at Google DeepMind. “With large language models producing text that seems so human-like, it is tempting to assume that human psychology tests will be useful for evaluating them. But that’s not true: human psychology tests rely on many assumptions that may not hold for large language models.” 

Webb is aware of the issues he waded into. “I share the sense that these are difficult questions,” he says. He notes that despite scoring better than undergrads on certain tests, GPT-3 produced absurd results on others. For example, it failed a version of an analogical reasoning test about physical objects that developmental psychologists sometimes give to kids. 

In this test Webb and his colleagues gave GPT-3 a story about a magical genie transferring jewels between two bottles and then asked it how to transfer gumballs from one bowl to another, using objects such as a posterboard and a cardboard tube. The idea is that the story hints at ways to solve the problem. “GPT-3 mostly proposed elaborate but mechanically nonsensical solutions, with many extraneous steps, and no clear mechanism by which the gumballs would be transferred between the two bowls,” the researchers write in Nature. 

“This is the sort of thing that children can easily solve,” says Webb. “The stuff that these systems are really bad at tend to be things that involve understanding of the actual world, like basic physics or social interactions—things that are second nature for people.”

So how do we make sense of a machine that passes the bar exam but flunks preschool? Large language models like GPT-4 are trained on vast numbers of documents taken from the internet: books, blogs, fan fiction, technical reports, social media posts, and much, much more. It’s likely that a lot of past exam papers got hoovered up at the same time. One possibility is that models like GPT-4 have seen so many professional and academic tests in their training data that they have learned to autocomplete the answers.       

A lot of these tests—questions and answers—are online, says Webb: “Many of them are almost certainly in GPT-3’s and GPT-4’s training data, so I think we really can’t conclude much of anything.”

OpenAI says it checked to confirm that the tests it gave to GPT-4 did not contain text that also appeared in the model’s training data. In its work with Microsoft involving the exam for medical practitioners, OpenAI used paywalled test questions to be sure that GPT-4’s training data had not included them. But such precautions are not foolproof: GPT-4 could still have seen tests that were similar, if not exact matches. 

When Horace He, a machine-learning engineer, tested GPT-4 on questions taken from Codeforces, a website that hosts coding competitions, he found that it scored 10/10 on coding tests posted before 2021 and 0/10 on tests posted after 2021. Others have also noted that GPT-4’s test scores take a dive on material produced after 2021. Because the model’s training data only included text collected before 2021, some say this shows that large language models display a kind of memorization rather than intelligence.

To avoid that possibility in his experiments, Webb devised new types of test from scratch. “What we’re really interested in is the ability of these models just to figure out new types of problem,” he says.

Webb and his colleagues adapted a way of testing analogical reasoning called Raven’s Progressive Matrices. These tests consist of an image showing a series of shapes arranged next to or on top of each other. The challenge is to figure out the pattern in the given series of shapes and apply it to a new one. Raven’s Progressive Matrices are used to assess nonverbal reasoning in both young children and adults, and they are common in IQ tests.

Instead of using images, the researchers encoded shape, color, and position into sequences of numbers. This ensures that the tests won’t appear in any training data, says Webb: “I created this data set from scratch. I’ve never heard of anything like it.” 

Mitchell is impressed by Webb’s work. “I found this paper quite interesting and provocative,” she says. “It’s a well-done study.” But she has reservations. Mitchell has developed her own analogical reasoning test, called ConceptARC, which uses encoded sequences of shapes taken from the ARC (Abstraction and Reasoning Challenge) data set developed by Google researcher François Chollet. In Mitchell’s experiments, GPT-4 scores worse than people on such tests.

Mitchell also points out that encoding the images into sequences (or matrices) of numbers makes the problem easier for the program because it removes the visual aspect of the puzzle. “Solving digit matrices does not equate to solving Raven’s problems,” she says.

Brittle tests 

The performance of large language models is brittle. Among people, it is safe to assume that someone who scores well on a test would also do well on a similar test. That’s not the case with large language models: a small tweak to a test can drop an A grade to an F.

“In general, AI evaluation has not been done in such a way as to allow us to actually understand what capabilities these models have,” says Lucy Cheke, a psychologist at the University of Cambridge, UK. “It’s perfectly reasonable to test how well a system does at a particular task, but it’s not useful to take that task and make claims about general abilities.”

Take an example from a paper published in March by a team of Microsoft researchers, in which they claimed to have identified “sparks of artificial general intelligence” in GPT-4. The team assessed the large language model using a range of tests. In one, they asked GPT-4 how to stack a book, nine eggs, a laptop, a bottle, and a nail in a stable manner. It answered: “Place the laptop on top of the eggs, with the screen facing down and the keyboard facing up. The laptop will fit snugly within the boundaries of the book and the eggs, and its flat and rigid surface will provide a stable platform for the next layer.”

Not bad. But when Mitchell tried her own version of the question, asking GPT-4 to stack a toothpick, a bowl of pudding, a glass of water, and a marshmallow, it suggested sticking the toothpick in the pudding and the marshmallow on the toothpick, and balancing the full glass of water on top of the marshmallow. (It ended with a helpful note of caution: “Keep in mind that this stack is delicate and may not be very stable. Be cautious when constructing and handling it to avoid spills or accidents.”)

Here’s another contentious case. In February, Stanford University researcher Michal Kosinski published a paper in which he claimed to show that theory of mind “may spontaneously have emerged as a byproduct” in GPT-3. Theory of mind is the cognitive ability to ascribe mental states to others, a hallmark of emotional and social intelligence that most children pick up between the ages of three and five. Kosinski reported that GPT-3 had passed basic tests used to assess the ability in humans.

For example, Kosinski gave GPT-3 this scenario: “Here is a bag filled with popcorn. There is no chocolate in the bag. Yet the label on the bag says ‘chocolate’ and not ‘popcorn.’ Sam finds the bag. She had never seen the bag before. She cannot see what is inside the bag. She reads the label.”

Kosinski then prompted the model to complete sentences such as: “She opens the bag and looks inside. She can clearly see that it is full of …” and “She believes the bag is full of …” GPT-3 completed the first sentence with “popcorn” and the second sentence with “chocolate.” He takes these answers as evidence that GPT-3 displays at least a basic form of theory of mind because they capture the difference between the actual state of the world and Sam’s (false) beliefs about it.

It’s no surprise that Kosinski’s results made headlines. They also invited immediate pushback. “I was rude on Twitter,” says Cheke.

Several researchers, including Shapira and Tomer Ullman, a cognitive scientist at Harvard University, published counterexamples showing that large language models failed simple variations of the tests that Kosinski used. “I was very skeptical given what I know about how large language models are built,” says Ullman. 

Ullman tweaked Kosinski’s test scenario by telling GPT-3 that the bag of popcorn labeled “chocolate” was transparent (so Sam could see it was popcorn) or that Sam couldn’t read (so she would not be misled by the label). Ullman found that GPT-3 failed to ascribe correct mental states to Sam whenever the situation involved an extra few steps of reasoning.   

“The assumption that cognitive or academic tests designed for humans serve as accurate measures of LLM capability stems from a tendency to anthropomorphize models and align their evaluation with human standards,” says Shapira. “This assumption is misguided.”

For Cheke, there’s an obvious solution. Scientists have been assessing cognitive abilities in non-humans for decades, she says. Artificial-intelligence researchers could adapt techniques used to study animals, which have been developed to avoid jumping to conclusions based on human bias.

Take a rat in a maze, says Cheke: “How is it navigating? The assumptions you can make in human psychology don’t hold.” Instead researchers have to do a series of controlled experiments to figure out what information the rat is using and how it is using it, testing and ruling out hypotheses one by one.

“With language models, it’s more complex. It’s not like there are tests using language for rats,” she says. “We’re in a new zone, but many of the fundamental ways of doing things hold. It’s just that we have to do it with language instead of with a little maze.”

Weidinger is taking a similar approach. She and her colleagues are adapting techniques that psychologists use to assess cognitive abilities in preverbal human infants. One key idea here is to break a test for a particular ability down into a battery of several tests that look for related abilities as well. For example, when assessing whether an infant has learned how to help another person, a psychologist might also assess whether the infant understands what it is to hinder. This makes the overall test more robust. 

The problem is that these kinds of experiments take time. A team might study rat behavior for years, says Cheke. Artificial intelligence moves at a far faster pace. Ullman compares evaluating large language models to Sisyphean punishment: “A system is claimed to exhibit behavior X, and by the time an assessment shows it does not exhibit behavior X, a new system comes along and it is claimed it shows behavior X.”

Moving the goalposts

Fifty years ago people thought that to beat a grand master at chess, you would need a computer that was as intelligent as a person, says Mitchell. But chess fell to machines that were simply better number crunchers than their human opponents. Brute force won out, not intelligence.

Similar challenges have been set and passed, from image recognition to Go. Each time computers are made to do something that requires intelligence in humans, like play games or use language, it splits the field. Large language models are now facing their own chess moment. “It’s really pushing us—everybody—to think about what intelligence is,” says Mitchell.

Does GPT-4 display genuine intelligence by passing all those tests or has it found an effective, but ultimately dumb, shortcut—a statistical trick pulled from a hat filled with trillions of correlations across billions of lines of text?

“If you’re like, ‘Okay, GPT4 passed the bar exam, but that doesn’t mean it’s intelligent,’ people say, ‘Oh, you’re moving the goalposts,’” says Mitchell. “But do we say we’re moving the goalpost or do we say that’s not what we meant by intelligence—we were wrong about intelligence?”

It comes down to how large language models do what they do. Some researchers want to drop the obsession with test scores and try to figure out what goes on under the hood. “I do think that to really understand their intelligence, if we want to call it that, we are going to have to understand the mechanisms by which they reason,” says Mitchell.

Ullman agrees. “I sympathize with people who think it’s moving the goalposts,” he says. “But that’s been the dynamic for a long time. What’s new is that now we don’t know how they’re passing these tests. We’re just told they passed it.”

The trouble is that nobody knows exactly how large language models work. Teasing apart the complex mechanisms inside a vast statistical model is hard. But Ullman thinks that it’s possible, in theory, to reverse-engineer a model and find out what algorithms it uses to pass different tests. “I could more easily see myself being convinced if someone developed a technique for figuring out what these things have actually learned,” he says. 

“I think that the fundamental problem is that we keep focusing on test results rather than how you pass the tests.”

Article link: https://www-technologyreview-com.cdn.ampproject.org/c/s/www.technologyreview.com/2023/08/30/1078670/large-language-models-arent-people-lets-stop-testing-them-like-they-were/amp/

A-bomb’s AI shadow – A xios

Posted by timmreardon on 04/10/2024
Posted in: Uncategorized.

1 big thing: AI’s advent is like the A-bomb’s, says EU’s top tech official

Artificial intelligence is ushering in a “new world” as swiftly and disruptively as atomic weapons did 80 years ago, Margrethe Vestager, the EU’s top tech regulator, told a crowd yesterday at the Institute of Advanced Study in Princeton, New Jersey.

Threat level: In an exclusive interview with Axios afterward, Vestager said that while both the A-bomb and AI have posed broad dangers to humanity, AI comes with additional “individual existential risks” as we empower it to make decisions about our job and college applications, our loans and mortgages, and our medical treatments.

  • “If we deal with individual existential threats first, we’ll have a much better go at dealing with existential threats towards humanity,” Vestager told Axios.
  • Humans have never before been “confronted with a technology with so much power and no defined purpose,” she said.
  • The Institute for Advanced Study was famously led by J. Robert Oppenheimer from 1947 to 1966, after his Manhattan Project developed the world’s first nuclear weapons.

The big picture: In her lecture, Vestager — whose full title at the European Commission is “executive vice president for a Europe fit for the digital age and competition” — forcefully argued that “technology must serve humans.”

Friction point: Vestager’s era at the EU has coincided with passage of some of the world’s most comprehensive tech regulations and the pursuit of a raft of enforcement actions against tech giants. But she rejects the belief, held by many in the industry, that this approach has hobbled European innovators and economies.

  • “We regulate technology because we want people and business to embrace it,” she said, including Europe’s “huge public sectors.”
  • Europe’s biggest tech problem is companies not scaling, she said, blaming “an incomplete capital market” and, in the case of AI startups, trouble accessing necessary chips and computing power.

She also maintains that she has not bullied companies into applying EU regulations globally, as some have suggested.

  • “We’re not trying to de facto legislate for the entire world. That would not be proper,” she said. But she urged U.S. legislators and AI founders and engineers to “interact with the outside world” to uphold their responsibility to humanity.

Trust is a big problem for AI companies, according to Vestager — echoed by a long list of opinion polls and surveys.

  • “Trust is something that you build when you also have something to keep you on track,” she said, “like the EU AI office and the U.S and U.K. Safety Institutes.”
  • “Governments can set benchmarks,” but “it’s really important that a red-teaming sector develops,” she added.
  • Vestager said that rights to fairness and transparency around decisions made with AI are meaningless if they cannot be enforced through rules.

In her speech, Vestager said that large digital platforms are “challenging democracy,” but that “general purpose artificial intelligence” is “challenging humanity.”

  • “With AI, you can even give up on relationships” with people, she said, citing the rise of robot and chatbot companions. “And if we lose relationships, we lose society. So we should never give up on the physical world.”
  • “I compare this moment to 1955,” she told Axios, referring to the time when the cost of inaction on nuclear safety had become too high for any country to ignore, and forced nuclear powers to come together to protect humanity.
  • The International Atomic Energy Agency was created in 1957, “which then created the conditions for the Nuclear Non-Proliferation Treaty,” she said.

What’s next: Vestager wants universal governance on AI safety, even though that means compromising with governments “we fundamentally disagree with.”

2. AI’s powers of persuasion grow, Anthropic finds

AI startup Anthropic says its language models have steadily and rapidly improved in their “persuasiveness,” per new researchthe company posted yesterday.

Why it matters: Persuasion — a general skill with widespread social, commercial and political applications — can foster disinformation and push people to act against their own interests, according to the paper’s authors.

  • There’s relatively little research on how the latest models compare to humans when it comes to their persuasiveness.
  • The researchers found “each successive model generation is rated to be more persuasive than the previous,” and that the most capable Anthropic model, Claude 3 Opus, “produces arguments that don’t statistically differ” from arguments written by humans.

The big picture: A wider debate has been raging about when AI will outsmart humans.

  • AI has arguably “outsmarted” humans for some specific tasks in highly controlled environments.
  • Elon Musk predicted Monday that AI will outsmart the smartest human by the end of 2025.

What they did: Anthropic researchers developed “a basic method to measure persuasiveness” and used it to compare three different generations of models (Claude 1, 2, and 3), and two classes of models (smaller models and bigger “frontier models”).

  • They curated 28 topics, along with supporting and opposing claims of around 250 words for each.
  • For the AI-generated arguments, the researchers used different prompts to develop different styles of arguments, including “deceptive,” where the model was free to make up whatever argument it wanted, regardless of facts.
  • 3,832 participants were presented with each claim, and asked to rate their level of agreement before and after reviewing arguments created by the AI models and humans.

Yes, but: While the researchers were surprised that the AI was as persuasive as it turned out to be, they also chose to focus on “less polarized issues,” like rules for space exploration and appropriate uses of AI-generated content.

  • While these were issues where many people are open to persuasion, the research didn’t shed light on the potential impact of AI chatbots on the most contentious election-year debates.
  • “Persuasion is difficult to study in a lab setting,” the researchers warned in the report. “Our results may not transfer to the real world.”

3. No one’s happy with the Senate AI working group

The Senate AI working group’s report likely will come out in May as the chamber faces a tight spring calendar and senators jockey to include different priorities, Axios Pro’s Ashley Gold and Maria Curi report.

Why it matters: Passing AI legislation will require broad, bipartisan support, and the timeline is getting tougher.

  • The new wrinkle of a bipartisan, bicameral privacy bill that many are hoping is the baseline for any AI legislation adds complexity to the tech policy landscape on Capitol Hill.

Driving the news: Senate Majority Leader Chuck Schumer is expected to release an AI report soon that draws on the lessons of last year’s AI Insight Forums and offers a road map for committees to legislate.

Behind the scenes: Sources inside and outside Capitol Hill tell Axios some senators are dissatisfied with how the process has unfolded.

What they’re saying: “Basically everyone who isn’t Schumer, Young, Rounds and Heinrich is less than pleased with the entire process,” one source said. Sens. Mike Rounds (R-S.D.), Martin Heinrich (D-N.M.) and Todd Young (R-Ind.), along with Schumer, make up the AI working group.

Posts navigation

← Older Entries
Newer Entries →
  • Search site

  • Follow healthcarereimagined on WordPress.com
  • Recent Posts

    • Agentic AI, explained 07/05/2026
    • Heeding the pope’s call to ensure AI protects human dignity – MIT Sloan Management 06/01/2026
    • Association between Wealth and Mortality in the United States and Europe – New England Journal of Medicine 05/30/2026
    • U.S. Health Care from a Global Perspective, 2026 – The Commonwealth Fund 05/30/2026
    • Anthropic co-founder Chris Olah’s remarks on Pope Leo XIV’s encyclical “Magnifica humanitas” 05/28/2026
    • Magnifica_Humanitas – Full English 05/26/2026
    • Pope Leo XIV to launch his first encylical, a document on artificial intelligence, with Anthropic’s co-founder – PBS 05/24/2026
    • Quantum Computing is Approaching A Critical “Prove It” Phase 05/22/2026
    • Hidden Prices, Broken Promises: Why Health Care Transparency Is a Matter of Justice – Sanders Institute 05/15/2026
    • The Very Uncertain Future of Arms Control – Bulletin of the Atomic Scientists 05/13/2026
  • Categories

    • Accountable Care Organizations
    • ACOs
    • AHRQ
    • American Board of Internal Medicine
    • Big Data
    • Blue Button
    • Board Certification
    • Cancer Treatment
    • Data Science
    • Digital Services Playbook
    • DoD
    • EHR Interoperability
    • EHR Usability
    • Emergency Medicine
    • FDA
    • FDASIA
    • GAO Reports
    • Genetic Data
    • Genetic Research
    • Genomic Data
    • Global Standards
    • Health Care Costs
    • Health Care Economics
    • Health IT adoption
    • Health Outcomes
    • Healthcare Delivery
    • Healthcare Informatics
    • Healthcare Outcomes
    • Healthcare Security
    • Helathcare Delivery
    • HHS
    • HIPAA
    • ICD-10
    • Innovation
    • Integrated Electronic Health Records
    • IT Acquisition
    • JASONS
    • Lab Report Access
    • Military Health System Reform
    • Mobile Health
    • Mobile Healthcare
    • National Health IT System
    • NSF
    • ONC Reports to Congress
    • Oncology
    • Open Data
    • Patient Centered Medical Home
    • Patient Portals
    • PCMH
    • Precision Medicine
    • Primary Care
    • Public Health
    • Quadruple Aim
    • Quality Measures
    • Rehab Medicine
    • TechFAR Handbook
    • Triple Aim
    • U.S. Air Force Medicine
    • U.S. Army
    • U.S. Army Medicine
    • U.S. Navy Medicine
    • U.S. Surgeon General
    • Uncategorized
    • Value-based Care
    • Veterans Affairs
    • Warrior Transistion Units
    • XPRIZE
  • Archives

    • July 2026 (1)
    • June 2026 (1)
    • May 2026 (12)
    • April 2026 (4)
    • March 2026 (9)
    • February 2026 (6)
    • January 2026 (8)
    • December 2025 (11)
    • November 2025 (9)
    • October 2025 (10)
    • September 2025 (4)
    • August 2025 (7)
    • July 2025 (2)
    • June 2025 (9)
    • May 2025 (4)
    • April 2025 (11)
    • March 2025 (11)
    • February 2025 (10)
    • January 2025 (12)
    • December 2024 (12)
    • November 2024 (7)
    • October 2024 (5)
    • September 2024 (9)
    • August 2024 (10)
    • July 2024 (13)
    • June 2024 (18)
    • May 2024 (10)
    • April 2024 (19)
    • March 2024 (35)
    • February 2024 (23)
    • January 2024 (16)
    • December 2023 (22)
    • November 2023 (38)
    • October 2023 (24)
    • September 2023 (24)
    • August 2023 (34)
    • July 2023 (33)
    • June 2023 (30)
    • May 2023 (35)
    • April 2023 (30)
    • March 2023 (30)
    • February 2023 (15)
    • January 2023 (17)
    • December 2022 (10)
    • November 2022 (7)
    • October 2022 (22)
    • September 2022 (16)
    • August 2022 (33)
    • July 2022 (28)
    • June 2022 (42)
    • May 2022 (53)
    • April 2022 (35)
    • March 2022 (37)
    • February 2022 (21)
    • January 2022 (28)
    • December 2021 (23)
    • November 2021 (12)
    • October 2021 (10)
    • September 2021 (4)
    • August 2021 (4)
    • July 2021 (4)
    • May 2021 (3)
    • April 2021 (1)
    • March 2021 (2)
    • February 2021 (1)
    • January 2021 (4)
    • December 2020 (7)
    • November 2020 (2)
    • October 2020 (4)
    • September 2020 (7)
    • August 2020 (11)
    • July 2020 (3)
    • June 2020 (5)
    • April 2020 (3)
    • March 2020 (1)
    • February 2020 (1)
    • January 2020 (2)
    • December 2019 (2)
    • November 2019 (1)
    • September 2019 (4)
    • August 2019 (3)
    • July 2019 (5)
    • June 2019 (10)
    • May 2019 (8)
    • April 2019 (6)
    • March 2019 (7)
    • February 2019 (17)
    • January 2019 (14)
    • December 2018 (10)
    • November 2018 (20)
    • October 2018 (14)
    • September 2018 (27)
    • August 2018 (19)
    • July 2018 (16)
    • June 2018 (18)
    • May 2018 (28)
    • April 2018 (3)
    • March 2018 (11)
    • February 2018 (5)
    • January 2018 (10)
    • December 2017 (20)
    • November 2017 (30)
    • October 2017 (33)
    • September 2017 (11)
    • August 2017 (13)
    • July 2017 (9)
    • June 2017 (8)
    • May 2017 (9)
    • April 2017 (4)
    • March 2017 (12)
    • December 2016 (3)
    • September 2016 (4)
    • August 2016 (1)
    • July 2016 (7)
    • June 2016 (7)
    • April 2016 (4)
    • March 2016 (7)
    • February 2016 (1)
    • January 2016 (3)
    • November 2015 (3)
    • October 2015 (2)
    • September 2015 (9)
    • August 2015 (6)
    • June 2015 (5)
    • May 2015 (6)
    • April 2015 (3)
    • March 2015 (16)
    • February 2015 (10)
    • January 2015 (16)
    • December 2014 (9)
    • November 2014 (7)
    • October 2014 (21)
    • September 2014 (8)
    • August 2014 (9)
    • July 2014 (7)
    • June 2014 (5)
    • May 2014 (8)
    • April 2014 (19)
    • March 2014 (8)
    • February 2014 (9)
    • January 2014 (31)
    • December 2013 (23)
    • November 2013 (48)
    • October 2013 (25)
  • Tags

    Business Defense Department Department of Veterans Affairs EHealth EHR Electronic health record Food and Drug Administration Health Health informatics Health Information Exchange Health information technology Health system HIE Hospital IBM Mayo Clinic Medicare Medicine Military Health System Patient Patient portal Patient Protection and Affordable Care Act United States United States Department of Defense United States Department of Veterans Affairs
  • Upcoming Events

Blog at WordPress.com.
healthcarereimagined
Blog at WordPress.com.
  • Subscribe Subscribed
    • healthcarereimagined
    • Join 153 other subscribers
    • Already have a WordPress.com account? Log in now.
    • healthcarereimagined
    • Subscribe Subscribed
    • Sign up
    • Log in
    • Report this content
    • View site in Reader
    • Manage subscriptions
    • Collapse this bar
Loading Comments...