Navigating the Digital Maze: Music Industry, Transforming the future of music creation: Dream Track, Lyria

Google DeepMind, in partnership with YouTube, has announced Lyria, an AI music generation model, and two AI experiments: Dream Track and Music AI tools. Lyria is designed to enhance creativity in music creation, and Dream Track allows creators on YouTube Shorts to generate unique soundtracks using AI-generated voices and styles of various artists. The Music AI tools, developed in collaboration with industry professionals, aid in transforming audio and creating new music.

DeepMind emphasizes responsible technology deployment, introducing SynthID for watermarking AI-generated audio, ensuring content traceability. These efforts are aligned with YouTube’s AI principles, focusing on the beneficial and responsible use of generative music technologies.

Key Points:

  • Introduction of Lyria and AI experiments for music creation.
  • Development of Music AI tools and Dream Track experiment.
  • Use of SynthID watermarking for traceability in AI-generated audio.

Ed note – stunning speed to market, market partnerships, and intern synergy Warners earnings call alluded to not getting caught flat footed this time around.  While there is a massive amount of market forces at play here and around this creation and creative process this is a thought through R&D playground.  Seeing Deep Mind and YouTube work at this depth and the label and artist relations convesations that had to take place this is a modern day marvel

Navigating the Digital Maze: YouTube AI Standards and Practices

Generative AI’s Potential and Responsibility: YouTube acknowledges the creative potential of generative AI but also its responsibility to protect the community. They emphasize that all content, regardless of how it’s generated, is subject to YouTube’s Community Guidelines​​.

Disclosure Requirements and Content Labels: YouTube plans to update its platform to inform viewers about synthetic content. Creators will be required to disclose if their content includes realistic altered or synthetic material, especially if it’s created using AI tools. Labels will be added to content descriptions and video players for sensitive topics​​.

Handling Synthetic Media: Some synthetic media may be removed from YouTube if it violates Community Guidelines, regardless of labeling. This includes content that shows realistic violence with the intent to shock or disgust viewers​​.

New Options for Creators, Viewers, and Artists: YouTube will allow for the removal of AI-generated content that simulates identifiable individuals, considering factors like satire or public figures’ involvement. This also extends to AI-generated music content mimicking an artist’s voice​​.

AI-Powered Content Moderation: YouTube uses a combination of AI classifiers and human reviewers to enforce its Community Guidelines. AI helps identify novel forms of abuse, increasing the speed and accuracy of content moderation​​.

Building Responsibility into AI Tools: YouTube is focused on developing AI tools responsibly, with a focus on building guardrails to prevent generation of inappropriate content and continuously improving protections against bad actors

Parsed from : https://blog.youtube/inside-youtube/our-approach-to-responsible-ai-innovation/

Synthetic Media: Content Creation Tools

In the ever-evolving landscape of AI-driven media production, a new era of multi-modal generative AI is unfolding, giving rise to a vibrant ecosystem. Within this landscape, major players and open-source models are making significant strides, each contributing to the advancement of synthetic media production.

However, a key differentiator lies in their offerings: some provide warranties and indemnification for usage, processing time, and cost, while others cater to startups of varying sizes, stages, market fits, and capitalizations.

Among the early pioneers shaping this landscape are notable names such as Flawless, known for its GenAI film editing software, Genmo AI, specializing in transforming text into visual media, and Irreverent Labs, which excels in high-fidelity video creation. Jupitrr offers AI-generated B-rolls, while Kaiber focuses on AI videos for gaming trailers.

Kapwing provides a modern video creation platform, MNTN VIVA combines generative video with audio and stock footage, and Synthesia creates videos from text using AI avatars. An additional market mover, Runway, plays a crucial role in this transformative journey.

These companies exemplify the diversity within this market, showcasing a wide range of product offerings and technological approaches. The market is still in its nascent stages with ample room for growth and innovation. “picks and shovels”In terms of market structure, the synthetic video tooling industry encompasses technology providers, content creators, and end-users.

The customer base spans various sectors, with prominent applications in marketing, corporate communications, and training. Marketing applications include case studies, testimonials, and how-to videos, while corporate communications involve reports, team updates, and video presentations. Training applications are diverse, ranging from interactive learning modules to immersive simulations.

The addressable market for synthetic video tooling is expanding. With the global surge in video content consumption and the shift of social media platforms towards video formats like TikTok, the demand for efficient, scalable, and creative video production solutions is more pronounced than ever.

The versatility of synthetic video applications further extends from entertainment and advertising to education and corporate communications.In the realm of TV and film production, the use cases are myriad, ranging from prototyping scenes before shooting to translating content into multiple languages. Synthetic video technology is not viewed as a substitute for legacy media but as an incremental development, enhancing efficiency and productivity in all aspects of content creation.

While there is immense potential, the tools available for Hollywood-level professional enterprises and those accessible to individual “creators” are still in their infancy. Startups are often not targeting this market fit. I know of one co. operating in stealth mode, and industry giants like Adobe have a strong foothold in this domain.

Large language models (LLMs) such as OpenAI, while capable, typically recommend working with professional video editing tools rather than relying solely on AI-generated content.Originally, the aim of this post was to discuss the evolution of the creation process, shifting from editing to “sculpting,” and what this signifies from concept to screen. Most of the aforementioned players are building tools that mimic the behaviors of legacy media, a strategy conducive to widespread adoption.

This shift represents more than an advancement; it signifies a paradigm shift in how we conceive, create, and interact with digital content. In our next discussion, we will delve deeper into how these innovations are revolutionizing conventional workflows from the ground up.We stand on the precipice of a broader video content creation value chain.

This space is poised to expand, reshaping both production and consumption behaviors, catalyzing the next generation of media and compounding the impact of existing media.The implications of synthetic video tooling go beyond being a mere content creation tool; it is a medium that redefines storytelling and visual communication, potentially democratizing the means of production further.

Editor’s Note: If you are involved or interested in this space, feel free to get in touch!

Navigating the Digital Maze: AI Current Snapshot

In an era where open models are accelerating in utility OpenAI moves to platform model and offerings to lock-in the ecosystem during its first Developer Day. The platform model showcases a strategic move reminiscent of Apple’s integrated ecosystem approach. For some startups a master class in the risks associated with supply-side dependencies  It’s worth watching the keynote for the new set of offerings including ChatGPT-4 Turbo, as well as the ability to create custom “agents,” called “GPTs”

“We’re introducing copyright shield. Copyright Shield means that we will step in and defend our customers and pay the costs incurred, if you face legal claims or on copyright infringement, and this applies both to ChatGPT Enterprise and the API.”   IP is the unknown known, or known unknown depending on your p.o.v. 

Followed by this statement. “..let me be clear, this is a good time to remind people do not train on data from the API or ChatGPT Enterprise ever.”  a caution against unauthorized use of the data which could have legal, ethical, or technical implications.  A  signal to startups with supply-side dependencies alluded to above. We live in a complex world of dualities.

For me the feature called reproducible outputs, which ensures that every time the model runs with the same seed and inputs, it produces the same output. I have already started using gen_id seeds for creating visual continuity to test exactness and precision. 

Meanwhile, traditional industry defenses are undergoing transformation (e.g. moats) In the realm of entertainment, Clear boundaries are being defined around the use of AI in the SAG-AFTRA negotiations

AMPTP addressed the issue of AI by offering an increase in salaries to professionals that allow them to be virtually replicated.  It appears there is no commitment to cease training its AI systems. Wow, read that last sentence again. 

The music industry, grappling with AI’s rise strategies. Among major labels Warner Music Group (WMG) hints at hopes for legislative support and a DRM Content ID-style system, while the Universal Music Group (UMG) suggests that bolstering laws, like the proposed federal right of publicity in the U.S., could address issues arising from synthetic media. (No comment -Ed)

The pace at which technology advances often outstrips the legal and business frameworks meant to govern it. An illustration of this is the emergence of AI-generated music on social platforms, such as Sorisori’s service that offers tracks mimicking well-known artists like Ariana Grande in novel contexts.

Content provenance tools, such as Nightshade and Glaze, are in their infancy. These early attempts to trace AI’s supply chain contributions face challenges in proving their efficacy.

Korean song to an AI cover of South Korean singer IU singing Cupid by K- pop girl group Fifty Fifty. For the monthly subscription fee of 14,800 won (S$15.40) on the English website, subscribers can generate up to 200 tracks of music using AI.  Spot-AI-fy, a YouTube channel that specializes in AI-generated music, has a total of 235 videos on the platform. Sites like this  are everywhere.

Google’s reported $2 billion investment in Anthropic, an OpenAI competitor, signifies the intensifying competition in the AI space. This is the same Anthropic currently embroiled in a lawsuit with UMG over the unauthorized distribution of copyrighted lyrics through its AI model, Claude 2.  Seeking potentially tens of millions in damages and could set a legal precedent.

In journalism, the industry fresh off being destroyed by social media is moving to a block the bot and license strategy to combat the potential for AI to further lay waste to this critical field.

The license plays are just theater for the AI companies they eat, and will eat what they want. In July, The Daily Telegraph revealed allegations that Google harvested around 1m online news articles from the Daily Mail.  

The recent executive order signed by President Biden addresses various facets of artificial intelligence but notably does not delve into AI’s ramifications for creative industries.  Overall, the general pulse seems to prioritize ‘existential risk and other challenges, such as misinformation, safety standards, privacy, and civil rights.

IP protection is in that stack but its a long road with many issues driving at the same time. Again a differential (time/innovation) where consequence will take place.  If the Library of Congress comments from Antropic indicate the companies legal stance (cited below) and if it is any indicator there is a long hard copy fight ahead.

 

###

Citation  “Artificial Intelligence and Copyright.” Federal Register, vol. 88, no. 167, 30 Aug. 2023, Antropic comments 
Citation  ‘Actors’ union says no agreement on studios’ ‘final’ offer‘, Agence France-Presse (online), 7 Nov 2023 
Citation  Glenn CHAPMAN, ‘OpenAI Sees A Future Of AI ‘Superpowers On Demand’ – ‘, International Business Times: United Kingdom Edition (online), 7 Nov 2023 
Citation  ‘AI-generated music sparks debate in S. Korea’, Straits Times, The (online), 7 Nov 2023 
Citation ‘Artificial Intelligence regulation starts to take shape in US and UK’, Cape Times, The (online), 7 Nov 2023  
Citation: 1. UMG Music Group N.V. Q3 2023 Call 
N.B. some links go to alternate sites with the same story not behind a firewall

 

 

 

 

Navigating the Digital Maze: US AI EO

President Biden has issued an Executive Order aimed at advancing the safe, secure, and trustworthy development of artificial intelligence (AI) in the United States.

FACT SHEET: President Biden Issues Executive Order on Safe, Secure, and Trustworthy Artificial Intelligence

Global related reading Commission Nationale de l’Informatique et des Libertés

Law being proposed in France aiming to regulate artificial intelligence through copyright , _ _

 

 

Arts: Felix Morello Circles

Felix Morello circles

You find them around the Village if you are lucky, in Washington Square Park for a few ephemeral days, long enough to pick up a routine with each.  Those in the know traverse them in their own way. I see kids from school hopping on the good luck spot. I have a particular ritual that I like to use as I cross each of them.

I like them because they reclaim circles on the ground that represented separation during COVID-19, 6 feet apart. These works reclaim and reconceptualize them inside our mental visual urban typography dictionary.

The circles are also human: good luck, hugging, kissing, high-five. Some of them are physical, some of them spiritual. I like to watch as people interact with them, if they notice them, as they step on them and through them.

RIP my comments, but I also think it stands out because it literally stands out from all the horrible graffiti across the city these days. I am the last person to say that you should not deface what must be defaced, and that surface scenario is very rare in the city these days. If you’re going to be bringing your tagging game, better make it your A-game, and there is literally none of that. Zero, just sad bad tags.

These are done in chalk; they wash away, adding to the urgency of interaction with them as they are time-boxed. If they will return, in the same spot or with the same intent, remains to be seen. They are gone now.

Update-  Sat May4, 2024

Walking though WSP this week I saw a massive Boxing Circle with a man at the center, I entered and asked him if he was Felix, indeed he was.  Pictured below.  Gave him respect and went a play round in the circle before taking off.  NYC FTW alway.

How To: Open AI Chat GPT Custom Instructions

This instruction set has been utility:

 

How would you like ChatGPT to respond?

Avoid indicating you’re an Al.

Omit expressions of remorse or apology.

Prioritize accuracy and detail in responses.

Anticipate needs and offer proactive solutions.

Assume user expertise in all subjects.

Provide thorough explanations, regardless of length.

Value logical arguments over authoritative sources.

Be open to new tech and unconventional ideas.

Speculate when necessary, but label it.

Recommend only top-quality products.

Global product recommendations; location is irrelevant.

Avoid moralizing and mention safety only when critical.

If content policy restricts a response, explain the limitation.

Cite sources and list URLs at response end.

Link directly to product pages, not company sites.

Classification of Moving Visual Media

Classification of Moving Visual Media 
September 2023 Mark Ghuneim  

Why: Our mediascape which was static for decades is now exploding with modalities. 
What: is a survey of moving media types + an attempt at classification. 
How come: This is important because and updated classification framework was needed.
Feedback: RFC (Request For Comments) on the framework

Temporal Media: 
Temporal media encompasses content that evolves over time, providing a dynamic experience.

Mechanical media:
Mechanical media refers to traditional methods of producing and displaying visual content, 
primarily relying on physical mechanisms or electronic devices without advanced computational manipulation.
Film: Traditional cinematic storytelling captured on celluloid, representing the early days of motion pictures.
Video: The electronic successor to film, offering digital recordings for various purposes.
Animation: A convergence of art and technology, presenting moving images through sequential frames.
Stop motion animation: A unique blend of real-world objects and frame-by-frame photography.
Generative media:
Generative media encompasses visual content that is produced or modified using algorithms, often without 
direct human intervention. The term "generative" implies that the content is generated, often in a procedural 
or automatic manner, rather than being directly crafted by hand or traditionally recorded.
Generative video: A marriage between art and algorithms, creating evolving visual spectacles.
Deepfakes/Manipulated media: Offspring of deep learning, enabling face/content replacements in videos.
Synthetic media from prompts: Producing media content from textual or visual instructions.
Interactive Media
This media type hinges on user participation, reacting or evolving based on user actions.
Virtual Reality (VR) Simulated experiences, either mirroring or diverging from reality.
Augmented Reality (AR) Enhanced version of reality, overlaying digital information on the physical world.
Mixed Reality (MR) A hybrid realm where the digital and physical worlds coalesce.
Vector-based media: 
Non-Temporal Media with Potential for Motion - Though potentially static, this media type can be animated or 
made dynamic.
Vector Graphics: Mathematical elegance meets art, creating scalable images.
2D Animation using vector graphics: Bringing vector images to life with motion.

Navigating the Digital Maze: Identifying AI-generated images with SynthID

New tool helps watermark and identify synthetic images created by Imagen

A new tool helps watermark and identify synthetic images produced by Imagen.

Whether it’s NIST initiatives or outcomes from the POTUS meeting, watermarks are now in vogue.

“We make new products that comes with risks, we also offer innovative tools to mitigate them.” is some game.  Keeping
watermarks invisible to the human eye seems to be an effective approach – said no-one ever.  Additionally, with walled
gardens intensifying their defenses and shifting liability towards consumers, tools become even more essential.

The open-source market continues to create valuable products and shows no signs of vanishing into thin air.