Tag: GPT-4o

  • GPT-4o Returns To ChatGPT After OpenAI Reverses Decision

    GPT-4o Returns To ChatGPT After OpenAI Reverses Decision

    OpenAI, the creators of ChatGPT, have recently reversed some of their previous decisions following a significant backlash from users upset by their changes. After launching the new GPT-5 model, OpenAI made several unexpected moves that sparked controversy.

    So, what exactly transpired? When OpenAI announced the release of GPT-5 on August 6th during a live stream, excitement was high. The company’s CEO, Sam Altman, introduced the latest model designed to enhance ChatGPT’s capabilities. Shortly after, access to older models was removed, leaving users no choice but to use the new version.

    However, the situation has since shifted. OpenAI has now allowed subscribers of ChatGPT Plus—those paying $20 a month—to access some legacy models again. Currently, only the GPT-4o model is available for this purpose.

    Many users had formed strong attachments to the previous AI personalities, especially with GPT-4o, which allowed for more personalized and detailed interactions. Previously, various models within ChatGPT catered to different needs—models 3 and 4o, for example, handled complex reasoning and coding tasks. But with GPT-5 aiming to integrate the best features of these older versions, OpenAI decided to eliminate the older options to streamline the user experience.

    This decision was met with immediate and intense criticism. Reddit threads filled with angry comments, and some users expressed how much they mourned the loss of their familiar models. One user even described feeling physically ill upon hearing they couldn’t access GPT-4o anymore, comparing the experience to losing a dear friend.

    OpenAI’s CEO participated in a Reddit “Ask Me Anything” session, where users voiced their disappointment over the lack of personality and individuality in GPT-5. One heartfelt comment likened GPT-5’s new personality to “wearing the skin of a dead friend,” highlighting how emotionally invested many users had become in their AI companions. Altman initially mentioned that the company was considering reintroducing legacy models, and after some limited testing, this option was made more widely available.

    Many appreciate GPT-5’s improved practicality, noting its enhanced multitasking and coding skills. Still, critics argue its writing abilities do not match those of GPT-4o or even GPT-5 itself. OpenAI’s goal was to develop a more versatile tool—not just a conversational assistant—but Altman later acknowledged that they underestimated how important some features of GPT-4o were for users. The new model is designed to reduce hallucinations, be less overly agreeable, and adopt a more professional tone, emphasizing safe responses and a balanced approach to sensitive topics.

    The rollout has not been entirely smooth. OpenAI has expanded the number of complex reasoning questions that Pro users can ask—up to 3,000 per week—showing an intent to incorporate user feedback and refine their offerings. Altman has even hinted at future adjustments, asking users during the AMA whether they’d prefer to focus solely on GPT-4o or consider the potential of GPT-4.5.

    Despite these efforts, the launch of GPT-5 has faced several challenges. Announcements have raised eyebrows among AI enthusiasts, and the platform remains somewhat unstable for non-Plus subscribers. The stability for paying users, however, appears to have improved significantly.

    While the transition to GPT-5 hasn’t been entirely seamless, many users continue exploring the different models and sharing their experiences. If you’re using ChatGPT, it’s worth trying out the new features and models—stay tuned for further updates as OpenAI continues to develop and refine its technology.

  • Midjourney Unveils New Image Model To Compete With GPT-4o

    Midjourney Unveils New Image Model To Compete With GPT-4o

    Initially regarded as one of the leading image generation models in the early AI landscape, MidJourney has seemingly been outpaced by more user-friendly and free tools such as Gemini, ChatGPT, and Bing. The recent upgrade to OpenAI’s GPT-4o model, which offers outstanding image generation and the ability to replicate real photos alongside producing flawless text, has added to MidJourney’s challenges. To remain competitive—especially in light of the Studio Ghibli-inspired AI art trend captivating the internet—MidJourney is introducing a revamped model with numerous enhancements.

    CEO David Holz shared insights on the new V7 model via MidJourney’s official Discord server and in a blog post. According to Holz, the new model is “more intelligent with text prompts” and generates images with “significantly increased” quality and “stunning textures.”

    The new model is designed to produce images much faster, approximately ten times quicker than the existing version, making it ideal for brainstorming and iterative processes. Users can activate the Conversational mode (accessible only on web) to recreate portions of an image without needing to rewrite the entire prompt or enter Edit mode. These images are of relatively lower quality and only cost half of what regular images do.

    One of the most exciting new features for our new V7 model is something we call “Draft Mode.” Draft mode is half the cost and 10 times the speed and it might be the best way to iterate on ideas ever. Try it with voice, think out loud and let our ideas flow like liquid dreams. pic.twitter.com/ANfTMC6Ej1

    — Midjourney (@midjourney) April 4, 2025

    When using the Discord app on a computer or mobile device, the Conversational mode is replaced by a Voice mode, allowing users to “think out loud” and have images generated fluidly, much like flowing dreams. This function is also integrated into the newly launched Draft mode.

    Furthermore, MidJourney V7 can operate in both Relax and Turbo modes, providing high-resolution images (as opposed to Draft mode), with the latter using double the credits for expedited image creation.

    At present, the new V7 model is missing some functionalities, and workflows reliant on upscaling, inpainting, and retexturing will revert to the previous V6.1 model. The V7 model also introduces Personalization, allowing users to save their preferences for image styles. This setup process takes around five minutes and involves choosing from a selection of 200 images to refine preferences.

    MidJourney is currently conducting a community-driven alpha testing phase for the new model and promises additional features in the upcoming 60 days. To test it out, users can type /settings into the chat box on Discord or the web platform, send the message, and then change the default model to V7 from the available settings.

  • I Tested Gemini’s Wild New Native Image Generation Feature

    I Tested Gemini’s Wild New Native Image Generation Feature

    The term ‘natively multimodal‘ has been making waves in the AI community for over a year, yet companies have only recently begun to fully harness the multimodal capabilities of their AI models. Google has now unveiled its latest “Gemini 2.0 Flash Experimental” model, which includes the ability to generate and edit images directly.

    You may be asking yourself, what’s the fuss about image generation? True, AI-generated images have been a feature in many popular chatbots like ChatGPT for some time. However, image generation in platforms like ChatGPT or Gemini typically involves sending prompts to specialized diffusion models such as Dall-E 3 or Imagen 3. These models are specifically trained to create images and function as add-ons to the primary AI model, rather than being integrated within it.

    In contrast, language-vision models like Gemini are inherently multimodal. They possess the unique capability to understand, create, and alter both text and images natively. Up until now, no tech company has provided this level of functionality to users. OpenAI introduced its own image generation feature with GPT-4o in 2024, but it was never made publicly available.

    With native image generation, you benefit from enhanced consistency since multimodal models are trained on extensive datasets that include various forms of content. This leads to a better grasp of concepts and a broader general knowledge base.

    In addition to generating images, you can effortlessly edit them using simple prompts. For instance, you can upload an image and request the model to add sunglasses, insert legible text, remove objects, and more. Unlike diffusion models, which regenerate the entire image each time you make a request, natively multimodal models ensure consistency across multiple edits.

    Native Image Generation with Gemini 2.0 Flash Experimental

    As of now, the native image generation feature is not available to the general public. The Gemini 2.0 Flash Experimental model with this capability can only be accessed through Google’s AI Studio (visit) at no cost.

    Having tried out the model on AI Studio, I found it to be a thrilling experience. To begin, I created a visual guide showcasing the consistency of Gemini’s image generation capabilities by asking it to illustrate the steps for making an omelet, generating an image for each step.

    The results were impressively consistent, with no noticeable glitches. Even small details, like the bowl, remained the same between images. The images can be downloaded in a resolution of 1024 x 680, allowing you to produce visual guides on a variety of topics.

    Next, I requested Gemini to create an aesthetically pleasing table and then to display the table from a central camera angle. It executed this task flawlessly. I then asked Gemini to add a PlayStation to the table and give me a closer look. Once again, it delivered beautifully, capturing the PS5’s reflection in a nearby mirror.

    Native Image Editing with Gemini 2.0 Flash Experimental

    To showcase Gemini’s image editing feature, I uploaded an image from my gallery and instructed Gemini 2.0 to remove a wine glass from the table. Afterward, I requested it to add mushrooms to a pizza and was impressed by the outcome. Then, I asked Gemini to include a croissant, and it delivered once again, demonstrating the full potential of AI image editing, thanks to Gemini’s native multimodal capabilities.

    Then, I uploaded a personal image and asked Gemini to add sunglasses, followed by text that read “Beebom” on my shirt. Both requests were executed adeptly.

    Lastly, I asked Gemini to colorize an image, which it executed beautifully. The end result was even more stunning than the original, free from glitches or distortions.

    colorizing images with Gemini

    There are countless possibilities you can explore with Gemini’s new multimodal capabilities. Google has done an impressive job integrating native image generation and editing, and I plan to use it extensively in the upcoming weeks to push its boundaries.

    Following the launch of Veo 2 for video generation and Imagen 3 for specialized image generation, it seems that Google is outpacing OpenAI in several areas beyond just text generation. It’ll be interesting to see how OpenAI responds to reclaim its position as a leader with ChatGPT.

    Arjun Sha

    Enthusiastic about Windows, ChromeOS, Android, and issues regarding security and privacy. I enjoy tackling everyday computing challenges.


  • Is GPT-4o Free?

    Is GPT-4o Free?

    OpenAI, a leading artificial intelligence research organization, recently unveiled its newest flagship model, GPT-4o, which has been making waves in the tech community.

    This advanced multimodal model boasts capabilities that go beyond its predecessors, offering enhanced performance in various aspects. One of the most significant questions surrounding GPT-4o is whether it is free or not.

    GPT-4o Pricing

    GPT-4o is not entirely free, but it does offer a free tier for users. According to OpenAI’s official pricing page, the model is available in different tiers, each with varying pricing structures.

    For instance, the 128K context model, which includes GPT-4o, is priced at $5.00 per 1 million tokens. This pricing model is based on the number of tokens used in requests to the model, with tokens being pieces of words, approximately 750 words per 1,000 tokens.

    Free Tier Availability

    GPT-4o is available in the free tier of ChatGPT starting today, and it is also accessible to subscribers of OpenAI’s premium ChatGPT Plus and Team plans.

    The free tier comes with usage limits, but Plus users will have a message limit that is up to 5x greater than free users. Team and Enterprise users will have even higher limits.

    Future Plans and Availability

    OpenAI has announced plans to roll out GPT-4o to ChatGPT Free users with usage limits today. Additionally, the company is set to launch a new Voice Mode with GPT-4o’s advanced capabilities in alpha in the coming weeks, with early access for Plus users as the feature rolls out more broadly.

    The pricing structure is based on the number of tokens used in requests to the model, with varying tiers available for different levels of usage. For those interested in accessing the advanced capabilities of GPT-4o, the free tier and premium plans offer a range of options to suit different needs and budgets.

  • GPT-4 vs GPT-4o: Which One Is Efficiency and Cost-Effective?

    GPT-4 vs GPT-4o: Which One Is Efficiency and Cost-Effective?

    In the rapidly evolving landscape of artificial intelligence, the quest for efficiency and cost-effectiveness is unending. Today, we’re comparing two heavyweights in the field: GPT-4 and its optimized version, GPT-4o.

    Speed and Efficiency

    GPT-4o has revolutionized the AI world with its speed. It’s not just faster than its predecessor, the GPT-4, but it’s also twice as fast as the GPT-4 Turbo. This significant speed difference makes GPT-4o a game-changer, leaving GPT-4 in the dust.

    In real-world applications, the speed of GPT-4o is evident. For instance, it can generate a CSV file in less than a minute, while GPT-4 takes nearly as long to generate the cities being used in the example. This speed makes GPT-4o much more usable for a variety of use cases, from data analysis to content generation.

    Cost-Effectiveness

    But GPT-4o isn’t just about speed. It’s also about cost. GPT-4o is 50% cheaper for developers to implement and has much higher rate limits. This makes it a more economical choice for developers looking to implement AI in their applications. Whether you’re a startup on a tight budget or a large corporation looking to cut costs, GPT-4o offers a cost-effective solution.

    Performance and Hallucinations

    Performance is a crucial factor in selecting the right model. While GPT-3.5-Turbo is effective for general-purpose tasks, it may encounter difficulties with highly complex queries.

    GPT-4o offers a balanced solution. It minimizes hallucinations more effectively than GPT-3.5-Turbo and approaches the reliability of GPT-4. This makes it suitable for applications requiring a blend of high performance and cost efficiency.

    GPT-4o is a faster, better successor to GPT-4. It combines the advanced functionality of GPT-4 with optimized efficiency, resulting in reduced costs. If you’re looking for high accuracy and reliability without the full expense of GPT-4, GPT-4o is the way to go.

    So, whether you’re a developer looking to implement AI in your app or a business looking to leverage AI for growth, GPT-4o is a worthy contender to consider. It’s not just about the cost and speed; it’s about getting the most out of your AI investment.

    Scalability and Integration

    The architecture of GPT-4o is designed with scalability in mind. It can handle an increasing number of requests without a drop in performance, making it ideal for businesses that are scaling up. Its compatibility with various programming languages and frameworks also means that integration into existing systems is smoother and less time-consuming.

    Environmental Impact

    Another aspect where GPT-4o shines is its reduced environmental footprint. Due to its efficiency, it requires less computational power, which translates to lower energy consumption. This is a step forward in making AI more sustainable and environmentally friendly.

    Future-Proofing

    With the pace of technological advancement, future-proofing is essential. GPT-4o’s design anticipates future developments, ensuring that it remains relevant and adaptable to upcoming changes in AI technology. This foresight protects investments and ensures that applications remain cutting-edge.

    GPT-4o stands out as a robust, efficient, and cost-effective AI model. Its impressive speed and performance, coupled with lower costs and an eye toward sustainability, make it an attractive option for anyone looking to harness the power of AI. As the AI landscape continues to grow, GPT-4o is poised to be at the forefront, driving innovation and efficiency.