Remove Duplicate Objects
Java coding interview problem for Collections: Remove Duplicate Objects.
Removing duplicate objects is one of the most common Java Collections interview problems.
Unlike primitive values:
int[]
removing duplicates from objects requires understanding:
- Object equality
equals()hashCode()- HashSet behavior
- Collection design
This problem teaches an important Java pattern:
Object List
↓
Define Equality Rules
↓
Detect Duplicate Objects
↓
Return Unique Objects
What are Duplicate Objects?
Two objects are considered duplicates when they represent the same logical data.
Example:
Employee objects:
Employee(101, "John", 90000)
Employee(101, "John", 90000)
Although they are two different objects in memory:
Object 1
Memory Address: 1001
Object 2
Memory Address: 2001
they contain the same information.
Therefore:
Duplicate
Object Reference vs Object Equality
Java has two different concepts:
Reference Equality
Uses:
==
Checks:
Are both references pointing to the same object?
Example:
employee1 == employee2
Logical Equality
Uses:
equals()
Checks:
Do objects contain the same values?
Example:
employee1.equals(employee2)
Example
Employee e1 =
new Employee(101,"John");
Employee e2 =
new Employee(101,"John");
Memory:
e1 → Object A
e2 → Object B
Using:
e1 == e2
Result:
false
because they are different objects.
Using:
e1.equals(e2)
Result:
true
if equals() is implemented correctly.
Why Remove Duplicate Objects is Asked in Interviews?
This problem tests:
1. Java Object Equality
Understanding:
equals()
hashCode()
2. Collections Knowledge
Understanding:
- HashSet
- LinkedHashSet
- TreeSet
- Streams
3. Object Design
Can you define:
When are two objects equal?
4. Real Application Thinking
Duplicate removal is common in:
- Data processing
- Database cleanup
- API responses
- Report generation
Real-World Applications
Database Records
Removing duplicate:
Customer records
Employee records
Transaction records
API Response Processing
Example:
API returns:
[
User1,
User2,
User1
]
Need:
[
User1,
User2
]
Data Migration
During migration:
Old System Data
+
New System Data
↓
Remove duplicates
Search Systems
Remove duplicate:
- Documents
- Keywords
- Results
Problem Statement
Given a list of Employee objects, remove duplicate employees.
Two employees are duplicates if:
id
and
name
are same
Employee Example
Input:
[
Employee(101,"John"),
Employee(102,"Alice"),
Employee(101,"John")
]
Output:
[
Employee(101,"John"),
Employee(102,"Alice")
]
Employee Class
class Employee {
private int id;
private String name;
private double salary;
public Employee(
int id,
String name,
double salary) {
this.id = id;
this.name = name;
this.salary = salary;
}
public int getId() {
return id;
}
public String getName() {
return name;
}
public double getSalary() {
return salary;
}
@Override
public String toString() {
return id +
" " +
name +
" " +
salary;
}
}
Duplicate Detection Concept
Java needs a rule:
How to identify duplicates?
Example:
Employee:
id = 101
name = John
salary = 90000
Another:
id = 101
name = John
salary = 100000
Are they duplicates?
Depends on business rules.
Possible rules:
Rule 1
Only ID matters:
same id = duplicate
Rule 2
ID + Name:
same id and name
Rule 3
All fields:
id + name + salary
equals() and hashCode() Relationship
Hash-based collections use:
HashSet
HashMap
They depend on:
equals()
+
hashCode()
equals() Rule
If:
a.equals(b)
is true,
then:
a.hashCode()
must equal:
b.hashCode()
Why hashCode is Needed?
HashSet internally works like:
Object
↓
hashCode()
↓
Bucket
↓
equals()
Example:
Adding:
Employee(101,"John")
HashSet calculates:
hashCode()
Finds bucket.
Adding duplicate:
Employee(101,"John")
HashSet:
- Checks hash code.
- Checks equals().
- Rejects duplicate.
HashSet Internal Working
Structure:
HashSet
|
HashMap
|
Buckets
|
Objects
Example:
Bucket 5
Employee(101,"John")
Duplicate:
Employee(101,"John")
goes to same bucket.
Then:
equals()
returns:
true
Duplicate removed.
Approach 1 — Brute Force Duplicate Removal
The simplest approach:
- Create new list.
- For every object:
- Check if already exists.
- Add only unique objects.
Algorithm
For every employee:
Is employee already present?
|
No
↓
Add employee
Java Program — Brute Force
import java.util.ArrayList;
import java.util.List;
public class RemoveDuplicateBruteForce {
public static List<Employee> removeDuplicates(
List<Employee> employees) {
List<Employee> result =
new ArrayList<>();
for(Employee employee :
employees) {
if(!result.contains(employee)) {
result.add(employee);
}
}
return result;
}
}
Problem with contains()
The contains() method internally uses:
equals()
Without proper equals():
Two identical objects:
Employee(101,"John")
Employee(101,"John")
will not be detected.
Complexity Analysis — Brute Force
For:
n objects
Each object checks previous objects.
Time:
O(n²)
Space:
O(n)
Advantages
- Simple.
- Easy to understand.
- No extra collection knowledge required.
Drawbacks
- Slow for large lists.
- Requires repeated comparisons.
- Not production friendly.
Approach 2 — Using HashSet
The optimized approach uses:
HashSet
because HashSet stores only unique elements.
Algorithm
- Create HashSet.
- Add all objects.
- Convert back to List.
Java Program — HashSet Approach
import java.util.*;
public class RemoveDuplicateHashSet {
public static List<Employee> removeDuplicates(
List<Employee> employees) {
Set<Employee> uniqueEmployees =
new HashSet<>(employees);
return new ArrayList<>(
uniqueEmployees
);
}
}
Requirement
Employee class must override:
equals()
hashCode()
Complete equals() and hashCode() Implementation
HashSet duplicate removal works only when objects correctly define:
equals()
hashCode()
Employee Class with equals() and hashCode()
import java.util.Objects;
class Employee {
private int id;
private String name;
private double salary;
public Employee(
int id,
String name,
double salary) {
this.id = id;
this.name = name;
this.salary = salary;
}
public int getId() {
return id;
}
public String getName() {
return name;
}
public double getSalary() {
return salary;
}
@Override
public boolean equals(
Object obj) {
if(this == obj) {
return true;
}
if(obj == null ||
getClass() != obj.getClass()) {
return false;
}
Employee employee =
(Employee) obj;
return id == employee.id
&&
Objects.equals(
name,
employee.name);
}
@Override
public int hashCode() {
return Objects.hash(
id,
name);
}
@Override
public String toString() {
return id +
" " +
name +
" " +
salary;
}
}
Why Include Same Fields in equals() and hashCode()?
Example:
Employee 1:
id = 101
name = John
Employee 2:
id = 101
name = John
Both generate:
same hashCode
same equals result
Therefore:
Duplicate
HashSet Duplicate Removal Example
import java.util.*;
public class HashSetExample {
public static void main(String[] args) {
List<Employee> employees =
new ArrayList<>();
employees.add(
new Employee(
101,
"John",
90000));
employees.add(
new Employee(
102,
"Alice",
120000));
employees.add(
new Employee(
101,
"John",
90000));
Set<Employee> unique =
new HashSet<>(
employees);
System.out.println(
unique);
}
}
Output
101 John 90000
102 Alice 120000
Duplicate:
101 John
removed.
Approach 3 — LinkedHashSet
A normal HashSet:
- Removes duplicates.
- Does not maintain order.
Example:
Input:
John
Alice
Bob
John
HashSet output:
Bob
John
Alice
Order is not guaranteed.
If insertion order matters:
Use:
LinkedHashSet
LinkedHashSet Example
Set<Employee> uniqueEmployees =
new LinkedHashSet<>(
employees);
Output:
John
Alice
Bob
LinkedHashSet Internal Working
Structure:
LinkedHashSet
|
Hash Table
+
Doubly Linked List
It provides:
Unique elements
+
Insertion order
Approach 4 — TreeSet for Sorted Unique Objects
TreeSet provides:
Unique objects
+
Sorted order
Example:
Sort employees by:
Employee ID
Need:
Comparable
or
Comparator
TreeSet Example
Set<Employee> employees =
new TreeSet<>(
Comparator.comparing(
Employee::getId
)
);
Output:
101 John
102 Alice
103 Bob
TreeSet Internal Working
TreeSet uses:
Red Black Tree
Operations:
Insert:
O(log n)
Search:
O(log n)
HashSet vs LinkedHashSet vs TreeSet
| Feature | HashSet | LinkedHashSet | TreeSet |
|---|---|---|---|
| Duplicate Removal | Yes | Yes | Yes |
| Order | No guarantee | Insertion order | Sorted order |
| Performance | O(1) | O(1) | O(log n) |
| Internal Structure | Hash Table | Hash + Linked List | Red Black Tree |
Approach 5 — Java Stream distinct()
Java Streams provide:
distinct()
to remove duplicates.
Example
List<Employee> uniqueEmployees =
employees.stream()
.distinct()
.toList();
How distinct() Works
Internally:
Stream
↓
HashSet
↓
Unique Elements
Requirement
Objects must implement:
equals()
hashCode()
Remove Duplicates Based on Specific Field
Real applications often define duplicates differently.
Example:
Employees:
101 John 90000
101 John 120000
Should they be duplicates?
If business rule says:
Same employee ID
then compare only:
id
Remove Duplicate Employees by ID
Using HashMap:
Map<Integer,Employee> map =
new HashMap<>();
for(Employee employee : employees) {
map.put(
employee.getId(),
employee);
}
List<Employee> result =
new ArrayList<>(
map.values());
Dry Run
Input:
101 John 90000
102 Alice 120000
101 John 95000
Map:
After first:
101 → John 90000
After second:
102 → Alice 120000
After third:
101 → John 95000
Existing value replaced.
Final:
101 John 95000
102 Alice 120000
Remove Duplicate Objects Using Email
Example:
User objects:
id
name
email
Duplicate rule:
Same email
Use:
Map<String,User> uniqueUsers =
new HashMap<>();
Key:
email
Group Objects Using HashMap
Another common variation:
Group employees by department.
Example:
IT:
John
Bob
HR:
Alice
Java:
Map<String,List<Employee>> grouped =
employees.stream()
.collect(
Collectors.groupingBy(
Employee::getDepartment
)
);
Comparable vs Comparator with Duplicate Removal
Sometimes duplicate removal and sorting are combined.
Example:
Remove duplicates
↓
Sort employees by salary
Using TreeSet:
Set<Employee> result =
new TreeSet<>(
Comparator
.comparing(Employee::getSalary)
);
Important Warning
Comparator defines equality in TreeSet.
Example:
Comparator.comparing(
Employee::getSalary
)
Two employees with same salary:
John 90000
Alice 90000
may be considered duplicates.
Better:
Comparator.comparing(
Employee::getSalary
)
.thenComparing(
Employee::getId
);
Primitive vs Object Collections
Java Collections work with objects.
Cannot:
HashSet<int>
Use:
HashSet<Integer>
Autoboxing:
int
↓
Integer
Common Interview Mistakes
Mistake 1
Forgetting hashCode().
Wrong:
override equals()
only.
Problem:
HashSet may not detect duplicates.
Mistake 2
Using == for objects.
Wrong:
employee1 == employee2
Correct:
employee1.equals(employee2)
Mistake 3
Incorrect equals/hashCode contract.
Rule:
If:
equals() == true
then:
hashCode() must be same
Mistake 4
Using HashSet when order matters.
Use:
LinkedHashSet
Edge Cases
| Case | Handling |
|---|---|
| Empty list | Return empty |
| Single object | Already unique |
| All duplicates | One object remains |
| Null objects | Handle separately |
| Different fields | Define business rule |
Interview Follow-up Questions
Q1. Remove duplicate Employee objects.
Q2. Why override equals() and hashCode()?
Q3. Difference between HashSet and TreeSet?
Q4. Remove duplicates without modifying original list.
Q5. Remove duplicates based on employee ID.
Q6. Remove duplicates using Streams.
Q7. How does HashSet identify duplicates?
Q8. Can TreeSet remove duplicates?
Related Java Collection Problems
- Count Word Frequency Using HashMap
- Sort Employees by Salary
- Group Employees by Department
- Find Duplicate Elements
- Two Sum Using HashMap
- Top K Frequent Elements
- First Non-Repeating Character
Key Takeaways
Removing duplicate objects requires understanding:
Object Equality
↓
equals()
↓
hashCode()
↓
Collections
Recommended approaches:
Fast unique removal:
HashSet
Preserve insertion order:
LinkedHashSet
Sorted unique objects:
TreeSet
Stream processing:
distinct()
Complexity:
HashSet:
Time: O(n)
Space: O(n)
Frequently Asked Interview Questions
Q1. Why does HashSet remove duplicates?
Because it uses:
hashCode()
+
equals()
Q2. Why override hashCode with equals?
To maintain HashMap/HashSet contract.
Q3. Which collection maintains insertion order?
LinkedHashSet
Q4. Which collection provides sorted unique objects?
TreeSet
Interview Tip
When asked:
"Remove duplicate objects in Java."
Explain:
- Define object equality.
- Override equals() and hashCode().
- Use HashSet for unique objects.
- Choose LinkedHashSet if order matters.
- Choose TreeSet if sorting is required.
For senior Java interviews, discuss:
- HashSet internals.
- Hash collision handling.
- Equality contract.
- Stream distinct implementation.
- Business-based duplicate rules.
This demonstrates strong understanding of Java Collections, OOP design, and production-level object handling.