# Create new column removing duplicate text

**URL:** <https://community.parabola.io/t/create-new-column-removing-duplicate-text/1986>\
**Category:** Ask a question\
**Tags:** Building-Flows\
**Created:** [August 25, 2021, 8:20am UTC](https://community.parabola.io/t/create-new-column-removing-duplicate-text/1986 "2021-08-25T08:20:41Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mateus\_Coelho](https://avatars.discourse-cdn.com/v4/letter/m/e9bcb4/32.png) [@Mateus\_Coelho](https://community.parabola.io/u/Mateus_Coelho)\
**Post date:** [August 25, 2021, 8:20am UTC](https://community.parabola.io/t/create-new-column-removing-duplicate-text/1986/1 "2021-08-25T08:20:41Z")

</div>

Hey Parabola experts,

I hope someone can help me. I have a table with two columns that you will notice duplicate text between both columns, as you can see in the screenshot below and highlight in red:

 ![Screenshot 2021-08-25 at 09.10.46](https://us1.discourse-cdn.com/flex020/uploads/parabola/original/2X/d/d66b183a409fffa7adb57d15a351bbe1e4289799.png)

Based on this scenario, I want to create a new column removing the duplicate text like the screenshot below

![Screenshot 2021-08-25 at 09.08.52](https://us1.discourse-cdn.com/flex020/uploads/parabola/original/2X/7/71cbe65539e463d29b2ed0a491e1b67b650c6cf3.png)

Does anyone have any ideas on how to achieve this?

---

<div class="post-metadata">

**Author:** ![daniel](https://sea2.discourse-cdn.com/flex020/user_avatar/community.parabola.io/daniel/32/732_2.png) [@daniel](https://community.parabola.io/u/daniel)\
**Post date:** [August 25, 2021, 4:11pm UTC](https://community.parabola.io/t/create-new-column-removing-duplicate-text/1986/2 "2021-08-25T16:11:56Z")

</div>

Hi Mateus,

Good question! The flow pictured below uses some of your example data to remove duplicate text between two columns:

 ![image](https://us1.discourse-cdn.com/flex020/uploads/parabola/original/2X/3/3aea2d43918d803730d39a61def34ae75e424481.jpeg)

Here’s a brief overview on how to set this up:

1. Use an **Insert row numbers** step after pulling in your data. This will help combine your newly formatted data with your original dataset for reference.
2. Use a **Split columns** step to create a new row for each word in a cell. Split the words using a space and dash delimiter.
3. Use a **Find overlap** step to remove words in your `accounts title (2)` column that also exist in your `accounts title` column.
4. Use a **Merge duplicate rows** step to merge the values in your newly de-duplicated column using a dash delimiter
5. Use a **Combine tables** step to join your data back together.

 ![image](https://us1.discourse-cdn.com/flex020/uploads/parabola/original/2X/a/a01109d562226efae625f88f6e85aeffb8169481.png)

Copy and paste the snippet below to duplicate the steps in this flow:  
**`parabola:cb:992b9b4d1b374bde9ab4b5bd8318be2a`**

Let me know if that helps!

---

<div class="post-metadata">

**Author:** ![Mateus\_Coelho](https://avatars.discourse-cdn.com/v4/letter/m/e9bcb4/32.png) [@Mateus\_Coelho](https://community.parabola.io/u/Mateus_Coelho)\
**Post date:** [August 26, 2021, 7:02pm UTC](https://community.parabola.io/t/create-new-column-removing-duplicate-text/1986/3 "2021-08-26T19:02:07Z")

</div>

> [@daniel](#):
>
> parabola:cb:992b9b4d1b374bde9ab4b5bd8318be2a

Thanks, Daniel!

It makes a lot of sense your flow, however, it didn’t work 100%

I copy the flow created but I don’t know why at the end it doesn’t show the information for all rows.

 ![Screenshot 2021-08-26 at 19.48.24](https://us1.discourse-cdn.com/flex020/uploads/parabola/original/2X/6/69cfdfb1650079ae68dee720b16945df76deea60.png)

Do you have any idea why it’s not working properly?

Below you will find a link for the data I’m using and a short video showing the steps used.

Data link: [Test 76 - Google Drive](https://docs.google.com/spreadsheets/d/e/2PACX-1vT4_lVZTo43LR8W5CN755JGgFfy3edw2Bt2ZaU9TzzoJeJ_bI2ZsQSpRxrBJct8nByjjgQ3m2Jt3JE-/pubhtml?gid=1307363078&single=true)

video: [Vidyard Recording](https://share.vidyard.com/watch/eSv31ipBZqK51kd4as7HwR)?

Thanks,  
Bruno

---

<div class="post-metadata">

**Author:** ![daniel](https://sea2.discourse-cdn.com/flex020/user_avatar/community.parabola.io/daniel/32/732_2.png) [@daniel](https://community.parabola.io/u/daniel)\
**Post date:** [August 26, 2021, 7:21pm UTC](https://community.parabola.io/t/create-new-column-removing-duplicate-text/1986/4 "2021-08-26T19:21:37Z")

</div>

Hi @Mateus_Coelho,

Can you try removing the **Find Overlap** in that snippet and replace it with a new one? Then, plug in the top **Split columns** step first and the bottom Split columns step back into the Find overlap step.

Additionally, it looks like some of your column values have triple dashes `---` and special characters with accents. Use a **Find and replace** step to replace all triple dashes `---` with a single dash `-` and all special characters with a standard vowel.

Let me now if that helps!

---

<div class="post-metadata">

**Author:** ![Mateus\_Coelho](https://avatars.discourse-cdn.com/v4/letter/m/e9bcb4/32.png) [@Mateus\_Coelho](https://community.parabola.io/u/Mateus_Coelho)\
**Post date:** [August 31, 2021, 4:59pm UTC](https://community.parabola.io/t/create-new-column-removing-duplicate-text/1986/5 "2021-08-31T16:59:57Z")

</div>

Thanks @daniel!

I had to do a massive data clean-up but the workflow worked in the end.
