Running Apache Spark Jobs on IBM watsonx.data
Learn how to build, submit, monitor, and troubleshoot Spark applications using the watsonx.data Spark Engine with real world examples and best practices on watsonx.data lakehouse.If you’re coming from...
View ArticleIBM watsonx.data Intelligence: Metadata Enrichment & Text-to-SQL
From Natural Language to SQL (Text2SQL): watsonx.data Intelligence with DB2IBM watsonx.data Intelligence helps you turn raw database tables into governed, semantically rich data assets that can be...
View ArticleLayers on Layers — How You Can Improve Your Recommendation Systems
Most recommendation systems fail for a simple reason: they expect a single score to do too much work.But, real systems aren’t that clean. Data arrives at different times. Quality varies wildly. User...
View ArticleProcessing Millions of Records on IBM watsonx
How we achieved 3× throughput processing millions of LLM API calls by switching from multithreading to async I/O on IBM watsonxWhen we set out to process over a million LLM API calls per day, we didn’t...
View ArticleBuilding Production-Ready Observability for vLLM
Monitor, trace, and visualize vLLM using OpenTelemetry, Prometheus, Grafana, and Jaeger for robust, scalable, and LLM operations.Picture this: You’ve just deployed a shiny new Large Language Model...
View ArticleAutoHRise: An AI-Powered Hiring Assistant with Agentic AI, Crew AI, and...
Understanding the power of the Agentic AI using Crew AI with the Watsonx AI and Discovery.AutoHRise an AI-powered recruitment agentic assistant, automates the hiring process, reducing time-to-hire and...
View ArticleAutoHRise: Resume Screening Using Crew AI, Watsonx AI and Discovery
Understanding the power of Agentic AI in automating resume screening using Crew AI, Watsonx AI, and semantic search with Watsonx Discovery.This is the second blog in our blog series on building agentic...
View Article⏳ Oracle INTERVALs: It’s Not Just Data, It’s Logic in Motion
In the world of replication, most fields travel quietly — just bits across wires.But INTERVAL? It carries meaning.It tells systems when to retry, how long to wait, what time window to honor.So when you...
View ArticleRule Output Settings within a Project in IBM Knowledge Catalog: Standardising...
Photo by DongGeun Lee on UnsplashIn today’s fast-paced, data-driven world, high-quality data — accurate, complete, and consistent — is foundational to everything from regulatory compliance and...
View ArticleVersion Control in IBM SPSS Collaboration and Deployment Services (CaDS) :...
Version Control in IBM SPSS Collaboration and Deployment Services (CaDS) : Save, Track, and Restore Your Modeler Streams EasilyIn this blog, I would like to demonstrate how IBM SPSS CaDS can be used to...
View ArticleEnabling SSL for Database in IBM SPSS CaDS on Liberty Server —...
Enabling SSL for Database in IBM SPSS CaDS on Liberty Server — Post-Installation GuideIf you’ve recently installed the SPSS Collaboration and Deployment Services (CaDS) on IBM Liberty and are wondering...
View ArticleThe IKEA of Data: How to Bring Modular Thinking to Your Data Architecture...
“Phew! Those dreaded (rather liked) 3-letter acronyms — IOT…”A few years ago, I found myself thinking about how messy IoT data could get — fast. I ended up comparing it to a supermarket: different...
View ArticleDvaita: The Dual Role of AI in Cybersecurity
Explore how generative AI like Dvaita can defend your systems — or destroy them — depending only on the prompt. Learn how to govern LLMs before they flip!Your Smartest Security Assistant. Your Most...
View ArticleSQL’s Midlife Crisis: From Manual Laborer to AI-Assisted Maestro
It was the summer of ’89.You were hunched over a beige CRT monitor. Cursor blinking. Someone nearby cursed — #!@% dropped a semicolon!You smirked and typed:SELECT * FROM employees WHERE salary >...
View ArticleGraceful External Termination: Handling Pod Deletions in Kubernetes Data...
Graceful External Termination: Handling Pod Deletions in Kubernetes Data Ingestion and Streaming JobsWhen running big-data pipelines in Kubernetes, especially streaming jobs, it’s easy to overlook how...
View ArticleGrouping of Rules: A Feature for Multiple Rule Run Management
Governing data better using bulk Data Quality Rule run in IBM Knowledge CatalogPhoto by Eric Prouzet on UnsplashGrouping of Data Quality Rules — yes, you got it right! But how does it work?Just like in...
View ArticlePushing the Boundaries of AI-based Lossy Compression
A CVPR EARTHVISION Data Challenge by Embed2ScaleModern compression methods redefine the way we handle and analyze satellite imagery. In this article, we introduce the 2025 CVPR EARTHVISION Data...
View ArticleNL2SQL with LangGraph Reflection Agent: Generating and Critiquing MySQL Queries
IntroductionTurning human language into SQL queries (NL2SQL) is incredibly useful, but how can we ensure the queries are correct? To achieve this, we need a technique or process that validates the...
View ArticleFine tuning AI models with InstructLab under IBM LSF
All the best for 2025! This blog looks back on a demo which I created for SC24 last November to show InstructLab workflows running on an IBM LSF cluster. Let’s begin with a bit of background. I’d like...
View ArticleServerless High Volume ETL data processing on Code Engine
By Santhosh Kumar Neerumalla, Niels Korschinsky & Christian HoeboerIntroductionThis blogpost describes how to manage and orchestrate high volume Extract-Transform-Load (ETL) loads using a...
View Article