Technology

Meta Unveils AI Models Capable of Generating Text and Images

Published

10 months ago

June 19, 2024

Meta has published five new research models on artificial intelligence (AI), some of which can detect AI-generated speech inside longer audio clips and others that can generate text and visuals.

Meta’s Fundamental AI Research (FAIR) team made the models publicly available on Tuesday, June 18, the firm announced in a news statement on Tuesday.

The announcement from Meta stated, “We hope that by making this research publicly available, we can inspire iterations and ultimately help advance AI in a responsible way.”

According to the announcement, one of the new models is called Chameleon, and it belongs to a family of mixed-modal models that can comprehend and produce both text and visuals. These models are able to produce a combination of text and images from input that contains both text and images. In the release, Meta hinted that this feature may be used to provide descriptions for photos or to build new scenes using text prompts and images.

Pretrained models for code completion were also provided on Tuesday. According to the press release, these models were trained using Meta’s novel multitoken prediction technique, which trains large language models (LLMs) to predict several upcoming words simultaneously rather than one at a time as in the past.

More control over AI music creation is available with the third new model, JASCO. According to the announcement, this new model can take in several inputs, such as beats or chords, for the purpose of generating music, instead of solely depending on text inputs. This feature enables the integration of audio and symbols into a single text-to-music production model.

According to a press release, another new model called AudioSeal has an audio watermarking technique that allows the localized identification of AI-generated speech, or the ability to identify specific AI-generated portions inside a larger audio clip. Additionally, this model is up to 485 times faster than earlier techniques at identifying speech produced by AI.

The goal of the sixth new AI research model that Meta’s FAIR team unveiled on Tuesday is to broaden the geographical and cultural diversity of text-to-image generating systems. In order to enhance assessments of text-to-image models, the company has made geographic disparities evaluation code and annotations available for this purpose.

Capital spending on AI and the metaverse development division Reality Labs is expected to reach $35 billion to $40 billion by the end of 2024, according to a metaverse financial report released in April. This represents an increase in spending of $5 billion over the company’s initial projections.

During the company’s quarterly earnings call on April 24, Meta CEO Mark Zuckerberg stated, “We’re building a number of different AI services, from our AI assistant to augmented reality apps and glasses, to APIs [application programming interfaces] that help creators engage their communities and that fans can interact with, to business AIs that we think every business on our platform will use.”

Up Next

Butterflies: World’s First Social Platform Connecting Humans and AI Launched

Don't Miss

Snap Introduces AI-Powered Augmented Reality Technologies

Kajal Chavan

Technology

Microsoft Expands Copilot Voice and Think Deeper

Published

2 months ago

February 25, 2025

Archana Suryawanshi

Microsoft Expands Copilot Voice and Think Deeper

Microsoft is taking a major step forward by offering unlimited access to Copilot Voice and Think Deeper, marking two years since the AI-powered Copilot was first integrated into Bing search. This update comes shortly after the tech giant revamped its Copilot Pro subscription and bundled advanced AI features into Microsoft 365.

What’s Changing?

Microsoft remains committed to its $20 per month Copilot Pro plan, ensuring that subscribers continue to enjoy premium benefits. According to the company, Copilot Pro users will receive:

Preferred access to the latest AI models during peak hours.
Early access to experimental AI features, with more updates expected soon.
Extended use of Copilot within popular Microsoft 365 apps like Word, Excel, and PowerPoint.

The Impact on Users

This move signals Microsoft’s dedication to enhancing AI-driven productivity tools. By expanding access to Copilot’s powerful features, users can expect improved efficiency, smarter assistance, and seamless integration across Microsoft’s ecosystem.

As AI technology continues to evolve, Microsoft is positioning itself at the forefront of innovation, ensuring both casual users and professionals can leverage the best AI tools available.

Stay tuned for further updates as Microsoft rolls out more enhancements to its AI offerings.

Technology

Google Launches Free AI Coding Tool for Individual Developers

Published

2 months ago

February 25, 2025

Archana Suryawanshi

Google Launches Free AI Coding Tool for Individual Developers

Google has introduced a free version of Gemini Code Assistant, its AI-powered coding assistant, for solo developers worldwide. The tool, previously available only to enterprise users, is now in public preview, making advanced AI-assisted coding accessible to students, freelancers, hobbyists, and startups.

More Features, Fewer Limits

Unlike competing tools such as GitHub Copilot, which limits free users to 2,000 code completions per month, Google is offering up to 180,000 code completions—a significantly higher cap designed to accommodate even the most active developers.

“Now anyone can easily learn, generate code snippets, debug, and modify applications without switching between multiple windows,” said Ryan J. Salva, Google’s senior director of product management.

AI-Powered Coding Assistance

Gemini Code Assist for individuals is powered by Google’s Gemini 2.0 AI model and offers:
Auto-completion of code while typing
Generation of entire code blocks based on prompts
Debugging assistance via an interactive chatbot

The tool integrates with popular developer environments like Visual Studio Code, GitHub, and JetBrains, supporting a wide range of programming languages. Developers can use natural language prompts, such as:
“Create an HTML form with fields for name, email, and message, plus a submit button.”

With support for 38 programming languages and a 128,000-token memory for processing complex prompts, Gemini Code Assist provides a robust AI-driven coding experience.

Enterprise Features Still Require a Subscription

While the free tier is generous, advanced features like productivity analytics, Google Cloud integrations, and custom AI tuning remain exclusive to paid Standard and Enterprise plans.

With this move, Google aims to compete more aggressively in the AI coding assistant market, offering developers a powerful and unrestricted alternative to existing tools.

Technology

Elon Musk Unveils Grok-3: A Game-Changing AI Chatbot to Rival ChatGPT

Published

2 months ago

February 19, 2025

Archana Suryawanshi

Elon Musk Unveils Grok-3: A Game-Changing AI Chatbot to Rival ChatGPT

Elon Musk’s artificial intelligence company xAI has unveiled its latest chatbot, Grok-3, which aims to compete with leading AI models such as OpenAI’s ChatGPT and China’s DeepSeek. Grok-3 is now available to Premium+ subscribers on Musk’s social media platform x (formerly Twitter) and is also available through xAI’s mobile app and the new SuperGrok subscription tier on Grok.com.

Advanced capabilities and performance

Grok-3 has ten times the computing power of its predecessor, Grok-2. Initial tests show that Grok-3 outperforms models from OpenAI, Google, and DeepSeek, particularly in areas such as math, science, and coding. The chatbot features advanced reasoning features capable of decomposing complex questions into manageable tasks. Users can interact with Grok-3 in two different ways: “Think,” which performs step-by-step reasoning, and “Big Brain,” which is designed for more difficult tasks.

Strategic Investments and Infrastructure

To support the development of Grok-3, xAI has made major investments in its supercomputer cluster, Colossus, which is currently the largest globally. This infrastructure underscores the company’s commitment to advancing AI technology and maintaining a competitive edge in the industry.

New Offerings and Future Plans

Along with Grok-3, xAI has also introduced a logic-based chatbot called DeepSearch, designed to enhance research, brainstorming, and data analysis tasks. This tool aims to provide users with more insightful and relevant information. Looking to the future, xAI plans to release Grok-2 as an open-source model, encouraging community participation and further development. Additionally, upcoming improvements for Grok-3 include a synthesized voice feature, which aims to improve user interaction and accessibility.

Market position and competition

The launch of Grok-3 positions xAI as a major competitor in the AI chatbot market, directly challenging established models from OpenAI and emerging competitors such as DeepSeek. While Grok-3’s performance claims are yet to be independently verified, early indications suggest it could have a significant impact on the AI landscape. xAI is actively seeking $10 billion in investment from major companies, demonstrating its strong belief in their technological advancements and market potential.