From Blurry to Brilliant: AI-Powered Tool Rescues Memories

Aug 18: Imagine finding a very old family photograph that has faded over time. Parts of the image are missing, the faces are blurry, and years spent inside a cardboard album have erased many details. Restoring such an image in a traditional manner would require painstaking manual work or specialised software trained on thousands of examples.

Media Pitch: How AI Is Helping Preserve Photographs and Bring Memories to Life — IITGN Research

Human faces contain details, such as the eyes, mouth, nose, hairline, and skin texture, that are necessary for recognition and interpretation. Accurate recovery of such details becomes difficult when an image is degraded. An image restoration system should not be limited to making an image look sharper; it should also preserve the facial structures and identities of the people in that image. Further, real-world images can suffer from multiple problems simultaneously, such as noise, missing sections, blurring, and compression artefacts. Think of it as solving a visual puzzle with many complex pieces! 

Several existing artificial intelligence -based image restoration methods depend heavily on supervised learning. They use large collections of paired images, in which each pair contains one clean and one artificially degraded image. While effective, this approach has limitations. Collecting suitable training datasets for every possible type of damage is expensive and often impossible. Additionally, as mentioned earlier, real-world images may exhibit multiple types of degradation simultaneously, creating conditions that differ significantly from those in the images used during training. 

What if a tool could automatically reconstruct the faces in images, even when it has not seen similarly damaged images before? Researchers from the Indian Institute of Technology Gandhinagar  and the Indian Institute of Technology BHU developed Generative Latent Inversion for Blind Face Restoration , an AI-powered face restoration framework that does not require paired training data. It uses a powerful face-generation model to imagine possible clean versions of a damaged image and gradually refines them until the generated face matches the available visual evidence. Their study was published in Pattern Recognition Letters. 

GenR performs StyleGAN3-based inversion. What is StyleGAN3? A powerful AI that generates incredibly realistic images by converting code into a picture. StyleGAN3-based inversion involves a reversal: determining exactly which code is needed to recreate a specific real-life image. This code provides access to parameters that help modify an image realistically. Imagine a grainy, pixelated photograph of a bank robber captured by an old security camera! GenR has the potential to reconstruct a crisp, high-definition face image of this robber that the police can run through modern facial recognition software.

GenR includes a three-stage optimisation process. The model begins with capturing the global structure and broad features such as identity and pose. Then, it refines the image by fine-tuning specific components such as the eyes, nose, and jawline. Finally, it focuses on fine textures and details, producing realistic skin and hair, as well as other subtle visual features. 

“A model may try really hard to refine a degraded image and invent details that do not belong to the original face in the process. This is a classic example of overfitting. Moving from coarse structure to fine details in a controlled manner is an attempt to minimise this risk of overfitting in GenR,” remarked Akbar Ali, a fourth-year PhD student in the Department of Computer Science and Engineering at IITGN. Mr Ali is the first author of this study. “This approach judges an image in a way similar to how a human eye would analyse it. It is an effort to keep the final output sharp, realistic, and perfectly recognisable.” 

The team evaluated the tool’s performance on four single and multiple image degradation tasks: 

  • De-noising comprised removing random visual noise; 

  • Upsampling consisted of converting low-resolution images into sharper versions; 

  • Inpainting, or fixing the missing pieces, intentionally damaged an image by covering up parts of it and then ‘healed’ it so the final result was realistic; and 

  • Deartifacting focused on improving image quality by addressing blur and pixelation. 

The findings suggested that GenR consistently produced visual improvements in images, demonstrating its effectiveness across a wide range of degradation scenarios. For single degradation tasks, the pipeline produced a clean image in about 30 seconds, making it one of the fastest among current leading methods! 

One limitation of this tool is the potential to generate clean, realistic images that do not match the original person when the original image is severely damaged. Further, adjusting many complex settings can cause the tool to memorise specific details too perfectly, leading to unrealistic results unless one uses strict regulations. Potential applications include preserving historical photographs and film footage and aiding forensic analysis with clearer facial reconstructions, among others. GenR may also improve image quality for platforms such as video conferencing and social media.

According to Dr Shamuganathan Raman,

“GenR represents a significant advancement in blind face restoration. It was interesting to see that this framework offers a flexible and data-efficient alternative to traditional supervised methods. Systems like GenR could play a crucial role in bringing images back to life. Future work may focus on models to improve robustness and generalisation.” Dr Raman is a Professor in the Departments of Computer Science & Engineering and Electrical Engineering. He is also the Head of the CSE department and the Principal Investigator at the Computer Vision, Imaging, and Graphics Lab. The team also had Dr Indra Deep Mastan, Assistant Professor, Department of Computer Science & Engineering, IIT BHU. 

Professor Raman was also part of a study on text-to-video generation. Published in the 2026 IEEE/CVF Winter Conference on Applications of Computer Vision proceedings, this work proposes a training-free approach to generating video from text. 

Imagine asking the AI to generate a video based on the prompt, ‘a girl runs in a park with a pink balloon in her hand and looks at a brown kitten sitting near a tree.’ It is possible for a conventional system to produce a good video, but with inconsistent details, such as the balloon turning from pink to blue or the tree disappearing, among other aspects. It is because when dealing with multiple elements or concepts in the prompt, AI may fail to pay enough attention to one of them over time, causing the corresponding object to get modified in the generated scene. As a potential solution to this problem, the researchers proposed a new approach, Video-ASTAR, which aims to keep the visual elements, like the pink balloon, tied to the words that describe them in the prompt. 

It is like keeping track of individual objects and their relationships across frames, rather than being overwhelmed by details, so that the generated video remains faithful to the original description. The team also had Dr Prajwal Singh, a postdoctoral fellow at the CVIG Lab, and Dr Kuldeep Kulkarni and Dr Harsh Rangwani, research scientists at Adobe Research, India.    

As we approach World Photography Day on August 19, innovations like GenR highlight the crucial role of AI-based tools in preserving visual memories. Such tools may support digital heritage preservation efforts nationally and globally. This work coincides with the vision of the Government of India’s IndiaAI Mission and Digital India program. It also resonates with UNESCO’s Memory of the World initiative, which emphasises that the world’s documentary heritage belongs to all and should be protected for all. 

On the other end of the spectrum, systems like Video-ASTAR highlight AI’s involvement in constructing visual memories for which no photograph was ever taken. Think of a grandfather describing to his grandchildren how he played hide-and-seek with his grandfather when he was five years old. There is no visual evidence for these children of how their great-great-grandfather dressed, walked, or talked. In such a case, the text-to-video tool could translate a detailed description into a visual approximation, giving the grandkids a way to take a peek into what existed only in their grandfather’s memory. In this manner, visual forms of memories could be passed down through generations. Hence, AI may provide assistance in both directions, helping us preserve the images we have inherited and also giving shape to memories for which no visual output existed. 

The authors acknowledged support from the Visvesvaraya PhD Scheme for study 1 and the Prime Minister Research Fellowship and the Jibaben Patel Chair in Artificial Intelligence for study 2. Further, study 2 was a part of Dr Prajwal Singh’s internship at Adobe.

Leave a Comment

Your email address will not be published. Required fields are marked *